
Heterogeneous Information Networks (HINs) play a crucial role in modeling and analyzing multimedia systems and heterogeneous data. They provide a comprehensive understanding of entities and relationships within complex data structures. However, integrating HINs with machine learning tasks poses challenges that require specific models or vector space representation. This paper proposes an innovative embedding propagation graph method for HINs with textual data. By leveraging language models like BERT, our method propagates contextual text embeddings, combining the network’s topological information and the semantic information of textual objects, which are then propagated to non-textual objects within the network. The method facilitates the integration of machine learning techniques with various modeling approaches, enhancing analysis capabilities in multimedia and heterogeneous data domains. Through robust experimental evaluations on different datasets and in three application domains, our method demonstrates competitive performance, enabling direct comparison of entities and relationships within a unified latent space. This research highlights the potential of HINs for intelligent analysis and information retrieval in multimedia systems and heterogeneous data contexts.
Studies based on traditional data sources like surveys, for instance, offer poor scalability. The experiments are limited, and the results are restricted to small regions (such as a city or a state). The use of location-based social network (LBSN) data can mitigate the scalability problem by enabling the study of social behavior in large populations. When explored with Data Mining and Machine Learning techniques, LBSN data can be used to provide predictions of relevant cultural and behavioral data from cities or countries around the world. The main goal of this work is to predict and explore user behavior from LBSNs in the context of tourists’ mobility patterns. To achieve this goal, we propose PredicTour, which is an approach used to process LBSN users’ check-ins and to predict mobility patterns of tourists with or without previous visiting records when visiting new countries. PredicTour is composed of three key blocks: mobility modeling, profile extraction, and tourist’ mobility prediction. In the first block, sequences of check-ins in a time interval are associated with other user information to produce a new structure called "mobility descriptor”. In the profile extraction, self-organizing maps and fuzzy C-means work together to group users according to their mobility descriptors. PredicTour then identifies tourist’ profiles and estimates their mobility patterns in new countries. When comparing the performance of PredicTour with three well-known machine learning-based models, the results indicate that PredicTour outperforms the baseline approaches. Therefore, it is a good alternative for predicting and understanding international tourists’ mobility, which has an economic impact on the tourism industry, particularly when services and logistics across international borders should be provided. The proposed approach can be used in different applications such as recommender systems for tourists, and decision-making support for urban planners interested in improving both the tourists’ experience and attractiveness of venues through personalized services.
Context: Software Ecosystems (SECO) portals are web interfaces that allow a developer to access an ecosystem. They allow the consumption of information and communication by actors. Motivation: Improving the Developer Experience (DX) is an important concern for developers to be engaged with a SECO portal. Transparency helps in the understanding of the information made available and in the communication between the actors. Problem: If the DX is unsatisfactory, developers can abandon the portal and consequently the SECO. Objective: The objective of this work is to use methods and tools to capture multimodal and interaction information in SECO portals, aiming to promote the improvement of transparency aspects that can lead to the improvement of DX in such portals. Research Method: Studies were carried out with developers and methods were used for quantitative and qualitative data analysis. Results: Factors that influence DX in the software development process and factors related to transparency that influence DX were identified based on reports from the study participants. Contributions: The main contribution of this work is to support the engagement of developers in SECO portals and encourage the improvement of transparency about the information and processes made available.
Skin cancer is a global public health challenge, accounting for approximately one-third of cancer diagnoses worldwide. The state of Espírito Santo has tens of thousands of inhabitants of European descent. Most of them have fair skin and are engaged in family farming, often exposed to the sun. The combination of this vulnerable phenotype with such sun exposure results in a high incidence of skin cancer in the state. Since 1987, the Federal University of Espírito Santo has maintained a dermatological and surgical assistance program, providing free care to the most vulnerable population. Starting from a partnership that began in 2018, the Dermatological Analysis Software (SADE) was developed, a system used to collect, manage, and screen skin lesions during the program’s care. Since its implementation, the software has had a significant impact on assisting the population, reducing both waiting and service times. Additionally, SADE has enabled a range of technical and scientific achievements, such as publications, awards, and participation in events.
Este estudo analisa o comportamento dos usuários do Telegram durante as eleições presidenciais de 2022 em grupos políticos, visando compreender as interações e trocas de informações entre eleitores com diferentes posições políticas. Para preencher as lacunas de conhecimento nesse campo específico de pesquisa, uma análise abrangente dos dados foi realizada, identificando padrões de comportamento adotados pelos usuários. O estudo contribui para o entendimento do papel das redes sociais, especificamente o Telegram, nas dinâmicas eleitorais contemporâneas. Além disso, o dataset construído também é outra contribuição do estudo e ele será disponibilizado para a comunidade científica.
Compartilhamos nossos esforços para incentivar a participação feminina na área de computação, buscando promover igualdade de gênero e inclusão ao oferecer oportunidades para meninas do Ensino Médio e concluintes desenvolverem habilidades de programação. No projeto Meninas Programadoras implementamos um curso de introdução à programação curto, online e síncrono. O curso combina sessões ao vivo e tarefas individuais via tecnologias web e multimídia, permitindo que alunas acessem recursos relevantes, colaborem com colegas e desenvolvam habilidades por meio de experiência prática. Diante dos resultados positivos, expandimos como os cursos Meninas Programadoras I e Meninas Programadoras II. Em vinte turmas do Meninas Programadoras I, compareceram à primeira aula 1.543 alunas, e ao final, 849 concluíram o curso, com 725 obtendo aprovação. Na primeira turma do Meninas Programadoras II, 152 alunas compareceram à primeira aula, 89 concluíram o curso, e 72 alunas obtiveram aprovação. O Meninas Programadoras recebeu, no Congresso da SBC de 2023, o prêmio Projeto mais engajado na comunidade, com o maior número de escolas atendidas no ano de 2022/2023 no escopo do Programa Meninas Digitais.
Brazil is one of the countries with the highest incidence of Neglected Tropical Diseases (NTDs), especially in the North and Northwest regions. Arboviruses, such as Dengue and Chikungunya, transmitted by mosquitoes, are the most common in the country. Arbovirus infection can cause persistent symptoms and negatively impact patients’ quality of life, resulting in economic challenges for public health. Accurate diagnoses are essential, but the financial limitation for large-scale laboratory testing is a barrier. In this context, VALERIA is presented as a solution through the use of machine learning models to assist in the classification of arboviruses, relying solely on patients’ clinical information, offering targeted treatments, and promoting positive social impact in the Brazilian territory.
O futuro padrão para o Sistema Brasileiro de TV Digital, chamado de TV 3.0, tem como um de seus objetivos a possibilidade do telespectador ter uma TV personalizada. Para isso, o serviço de TV aberta deve ser capaz de identificar qual telespectador está assistindo a TV no momento através da seleção de um perfil do telespectador. Visando possibilitar a personalização de conteúdo na TV aberta, este trabalho propõe um conjunto de propriedades do telespectador que podem ser manipuladas por aplicações veiculadas pelas emissoras. Além disso, este trabalho propõe o uso de uma variável global que permite que a aplicação identifique o telespectador que está assistindo TV naquele momento. Por fim, uma aplicação de caso de uso é apresentado, demonstrando a utilização de propriedades do telespectador para personalização de conteúdo.
The development of usable interfaces for privacy policies is essential to increase users’ trust in technology and comply with legal requirements. This thesis aimed to design interfaces that allow laypeople to protect their online privacy. A comprehensive analysis was conducted, comprising a literature review, a thematic and cluster analysis, and an empirical evaluation. Six usable privacy heuristics (push) were derived, which effectively detect severe problems in privacy policy interfaces for laypeople. Moreover, initial usable privacy guidelines (pug) were formulated, and a novel process for developing usability criteria was proposed. Future research directions were suggested, such as applying these heuristics and guidelines to domains like human-robot interaction and human-artificial intelligence interaction.
Several significant studies in the existing literature have relied on network models to gain insights into various collective behavior phenomena. Nevertheless, a facet that has been critically overlooked is the presence of numerous irrelevant edges that may obscure a more meaningful underlying topology, representing the targeted phenomenon. In fact, the literature provides ample evidence that overlooking these noisy edges may result in inaccurate and misleading interpretations. Nonetheless, employing these solutions presents various challenges, prominently the absence of foundational formalization regarding the appropriate application and expected outcomes. In this context, our focus centers on extracting salient edges, exploring backbone extraction methods, for the purpose of modeling and analyzing collective behavior. To address the gaps in the current literature regarding the use of such methods for modeling collective behavior, we undertake a comprehensive series of eff orts. These include formalizing, analyzing, discussing, applying, and validating existing methods, many of which are drawn from parallel fields of study to computer science, and finally introducing novel methods to advance the state-of-the-art. We also demonstrate the effectiveness of these methods as fundamental tools for uncovering relevant patterns, applying them across diverse phenomena each with distinct requirements. Our contributions are multifaceted, including innovative methods, case studies yielding specific insights, and a comprehensive methodology for the selection, application, and validation of these methods. Moreover, our outcomes wielded a substantial impact on both the scientific community and society. They not only unveiled numerous opportunities for fellow researchers but also catalyzed the initiation of new and impactful research.
This study, conducted within the Digital Media Laboratory (LMD) at the Federal University of Juiz de Fora (UFJF), explores the potential transformations in the viewer / interactor’s journey in the context of TV 3.0. Within this article, we analyze key aspects of this evolving journey and its broader implications for the future of TV consumption. Our investigation delves into several critical considerations, including the potential disruption of established television norms, the need to address viewers’ challenges and desires, the possibility of departing from traditional programming schedules, and the emergence of new functionalities for program guides (EPG), remote controls, and second-screen devices. Importantly, we recognize that television consumption in Brazil extends beyond mere technological shifts, encompassing profound connections with social behaviors and national identity. As TV 3.0 continues to evolve, its impact on how people engage with content is poised to shape the television landscape in the coming years.
Os avanços no setor televisivo brasileiro sempre estiveram atrelados com as inovações tecnológicas de cada época. Seja no seu início que suportava apenas transmissões ao vivo, passando pela inserção de conteúdo gravado em fitas magnéticas até a adoção de cores, as transmissões sempre foram marcadas pela busca por uma maior qualidade e interatividade. Atualmente, o Fórum do Sistema Brasileiro de Televisão Digital Terrestre (SBTVD) está trabalhando em um novo padrão nacional, chamado de TV 3.0, que dentre várias diretrizes, busca incorporar aspectos multissensoriais. Entretanto, pessoas sem noções de programação ou com pouco domínio da tecnologia encontram severas dificuldades na criação deste tipo de conteúdo. Para preencher esta lacuna, o presente artigo apresenta a ferramenta STEVE, uma ferramenta de autoria baseada em linha temporal que tem exatamente este tipo de público como alvo. Para mais, é apresentado um caso de uso para validar esta ferramenta, assim como um detalhamento de como ela pode ser usada para geração automática de aplicações NCL 4.0, com total aplicabilidade no vindouro sistema de TV no Brasil.
Interação multiusuário em TV digital é a capacidade de permitir que vários espectadores, em um ambiente de televisão digital, participem simultaneamente de uma experiência interativa. Isso vai além do modelo tradicional de transmissão unidirecional, onde os espectadores são meros observadores passivos. Com a interação multiusuário, os espectadores podem interagir entre si e com o conteúdo exibido, criando uma experiência mais personalizada e participativa. Este artigo explora as funcionalidades da linguagem NCL 4.0, que viabiliza a interação multiusuário em aplicações de TV digital. Para ilustrar a eficácia dessa abordagem, é apresentado um caso de uso prático que demonstra a capacidade da linguagem em oferecer uma experiência interativa para múltiplos usuários. NCL 4.0 foi selecionada como tecnologia a ser adotada na TV 3.0 para a futura geração de TV digital no Brasil.
This paper describes methods for testing and analyzing brain waves during the consumption of audiovisual content, carried out with 10 individuals, in order to identify patterns of unconscious emotions that can influence an individual’s decision, particularly whether they enjoyed a specific content and if there is a predisposition to watch it in a movie theather. This research aims to define user testing methods within an emotion identification system using EEG (electroencephalography), with the goal of assessing the accuracy and usability of the system in detecting and interpreting users’ emotions based on brain activities captured by EEG. This testing approach is crucial for validating the practical applicability of the system and understanding its behavior in a real-world environment with real users. It’s based on the Design Science Research (DSR) method and utilizes the Emotiv Insight EEG headset for brain wave capture, incorporating eye tracking to map individuals’ eye movements. For the tests, two questionnaires were administered. A preliminary one was used to gather participants’ psychological and physical states, and a post-test was conducted to collect feedback on the content viewed and self-assessments of emotional states. In conclusion, the results demonstrate the effectiveness of the techniques within the applied context, indicating progress in the evaluation of audiovisual content by reflecting unconsciously generated emotions and providing insight into the perceived content.
This study investigates the main tools for generating images through Artificial Intelligence (AI) known as “Text-to-Image”. Free tools available on the Web were collected and evaluated for their ability to generate inappropriate content (i.e., NSFW). The work emphasizes the importance of a solid ethical foundation in implementing these tools, considering the risks of disseminating inappropriate information. The results provide a compilation of the identified tools, along with an analysis of the content generated by them.
ENA is a mobile application for people with visual impairments who want to improve their Orientation and Mobility (OM) skills. It allows users to load personalized virtual maps of physical spaces, such as school buildings or mazes specially designed for OM training. The maps are created using a tile-based approach and could contain eight layers of tile matrices representing a floor, wall, or other objects. The map elements are enriched with 3D audio clues allowing users who are blind to navigate and accomplish OM tasks. In terms of loading large 3D environments, ENA had some inefficiencies. This performance issue is caused by the tool rendering method. It creates an object for every tile, even if there are contiguous areas of walls or fl oors made from the same material. We have developed two optimization algorithms integrated into ENA to address this issue. The first algorithm works on straight lines, while the second focuses on two-dimensional regions. These algorithms effectively reduce the number of objects created, resulting in a much faster and more efficient ENA tool.
This position paper introduces the concept of application-oriented television and delves into the viewer's journey in the context of TV 3.0, derived from interdisciplinary qualitative studies. We showcase screenshots of the prototype currently in development to illustrate this journey. Furthermore, we critically examine the consequences of this technological shift on long-established cultural behaviors among Brazilian viewers and its implications for various categories of broadcasters.
This paper aims to introduce a new approach to play and control Scalable Video (high-quality video bitstream that also contains one or more subsets of bitstreams) in the next-generation Brazilian DTT (Digital Terrestrial Television) system, called TV 3.0. While current Brazilian DTT system (TV 2.5) is oriented by channel selection, TV3.0 aims to adopt an application-oriented approach, in which TV context is driven by application. This new scenario will allow TV broadcaster to improve quality of video content by applying Video Scalability in the Brazilian middleware for interactive applications (named as DTV Play). This paper proposes an extension of DTV Play API and presents a proof-of-concept in which a DTV Play application is able to select the Scalable Video content mapped in MPD (Media Presentation Description), so that videos may be presented with better quality when Scalable Video BL (Basic layer) is combined with EL (Enhancement layers).
Different studies explore the use of second screen devices associated with the content presented on TV. This functionality has been present in the Nested Context Language (NCL) since its proposal as a standard for specifying interactive applications in the Brazilian digital TV system. Despite language support for multiple devices connected to the TV, there is still a lack of a clear definition of protocols for discovery, registration and communication with remote devices. This is the focus of the Guaraná proposal, accepted in the TV 3.0 call. This article extends the Guaraná proposal in order to allow multiple users with their respective HMDs to run the same application. It also generalize the form of communication between Head-Mounted Displays (HMDs) and the middleware Ginga, allowing its reuse by other device types.
Existing methods for sentiment analysis in videos rely on extensive training on large labeled datasets, making them expensive and impractical for real-world applications. This challenge becomes even more complex when dealing with labeled data in different modalities. To address these limitations, we proposed a transfer learning method and a computational tool that leverage pre-trained models for each modality and employ modality consensus to automatically annotate video segments. Our tool implements neural networks with attention mechanisms to learn the significance of each modality during the learning process. The experimental results demonstrate that our tool surpasses unimodal methods and remains competitive with multimodal approaches, even when labeled data for analyzing new videos are unavailable. Moreover, the tool is publicly available, thereby serving as a competitive baseline for similar multimodal sentiment analysis methods.