In the fast-moving field of Natural Language Processing (NLP), making lexicons is still a method for many text analysis applications. This process of generating lexicons has traditionally used techniques such as semantic matches, word embeddings, and tools like EMPATH. With the arrival of Large Language Models (LLMs) including GPT-3.5, GPT-4 and Mistral 7b 0.1, we have new ways to create lexicons. This study takes a close look at how these older methods stack up against the newer options brought by LLMs. We carried out a detailed analysis, looking at how well different methods could create lexicons, focusing on their precision, scalability, and concluding on how efficiently they can be used in real-world settings. By using standard NLP tasks like document classification, emotion classification and sentiment analysis, this research prove itself on a variety of datasets to test how well the lexicons worked. This discovery, along with others from our study, aims to help professionals and researchers find the best approaches to lexicon creation today, setting the stage for more research in the NLP field.
Collision avoidance in real-world environments and computer games with multiple moving obstacles remains a challenging problem due to the dynamic and unpredictable nature of obstacle motion. In this study, we propose a novel method that dynamically generates an artificial dangerous field based on obstacle position information detected by ray-based sensors and uses it for learning to avoid obstacles. Unlike traditional approaches that rely on raw positional inputs or assume global knowledge, our method computes a localized dangerous field from partial sensor data, enhancing real-time adaptability. We implement this method within a deep reinforcement learning framework and evaluate its effectiveness in a simulated environment where moving obstacles are continuously generated with randomized trajectories. Experimental results demonstrate that the proposed method consistently outperforms a baseline approach that directly uses obstacle position vectors, yielding higher average rewards in the later stages of training. This result suggests that the proposed dangerous field-based approach is effective for efficient and adaptive collision avoidance in dynamic environments.
In recent years, the spread of fake news has become a problem due to the popularization of social networking services (SNS). Fact-checking exists as a countermeasure, but it is pointed out that it may not function perfectly in SNS environments where the echo chamber phenomenon has progressed. In this study, we constructed a SNS model in which facts and opinions are mixed and simulated the effect of fact checking. The results showed that the number of views increased linearly with each delay in fact-checking, regardless of the degree of progression of the echo chamber phenomenon.
In sports and traditional arts, skilled people have sensory knowledge obtained by repetitive training, called as implicit knowledge. Implicit knowledge is difficult to transfer systematically, therefore, efficient transfer is expected to improve competitiveness in sports and resolve the lack of successors in traditional arts. In addition, instructor evaluation is important when acquiring skills. However, there are limited opportunities to get advice from instructors.Therefore, in this research, our aim is to develop a system that reproduces instructor evaluations using acceleration that can be acquired with a smartphone. Yamanaka et al. developed a martial arts demonstration evaluation system using acceleration data. However, the entire movement was input regardless of the evaluation items and the points of focus in the actual evaluation were different from the points evaluated by the system. In this study, in order to reproduce actual evaluations, we proposed a machine learning model using only the focus points for each evaluation item. In experiments, the accuracy improved when the entire movement data was changed to the data of only the point of focus. We also obtained results suggesting that areas other than the focus area do not contribute to the evaluation.
This study deals with transfer reinforcement learning, i.e., reinforcement learning using knowledge acquired in one learning process (source task) in other learning processes (target tasks). The knowledge improves performance in the early stages of learning, but it may sometimes be detrimental, which is called negative transfer. To avoid the negative transfer, we propose a method in which the agent itself measures an effect of the transferred knowledge on the current learning process, and if it is harmful, it discards the knowledge during learning. We conducted experiments in a multi-player game environment with independently learning agents, in which the agents first learned their policies in one map of the environment and after that they learned in other, randomly generated maps. The result shows that the proposal mitigated negative transfer more successfully than existing methods.
Spreading of smartphone, wearable devices, and IoT systems has led to an increasing amount of time-series data. Various types of time-series data are being sampled today, and it needs to be utilized efficiently. However, time-series data has several characteristics that differ according to sampled target, such as periodicity, trends, and complexity. The characteristics make analyzing methods harder to apply versatilely. Therefore, lots of studies focus on one target and propose a corresponding method. In this research, we focus on "momentary large value" and "temporary stopping" in such as martial arts demonstration, CNC control signal, and steep step response signal. In this paper, we propose a method which detects breaks of periods and divides data using spectrogram. The method assumes that the data has one peak and one stasis for each period and detects breaks of periods from multiple criteria by detecting each. In an experiment, we applied the method for 3-axes accelerometer data of Taekwondo demonstration sampled by smartphones. We confirmed the method could detect the breaks correctly.
In the rapidly changing world of smart technology, searching for documents has become more challenging due to the rise of advanced language models. These models sometimes face difficulties, like providing inaccurate information, commonly known as "hallucination." This research focuses on addressing this issue through Retrieval-Augmented Generation (RAG), a technique that guides models to give accurate responses based on real facts. To overcome scalability issues, the study explores connecting user queries with sophisticated language models such as BERT and Orca2, using an innovative query optimization process. The study unfolds in three scenarios: first, without RAG, second, without additional assistance, and finally, with extra help. Choosing the compact yet efficient Orca2 7B model demonstrates a smart use of computing resources. The empirical results, generated when we asked questions regarding schizophrenia, indicate a significant improvement in the initial language model’s performance under RAG, particularly when assisted with prompts augmenters. Consistency in document retrieval across different encodings highlights the effectiveness of using language model-generated queries. The introduction of UMAP for BERT further simplifies document retrieval while maintaining strong results.
POS data analysis is a technique for analyzing consumer purchasing behavior. One of the methods is to classify stores by extracting order trends from each receipt using non-negative matrix factorization(NMF). However, the problem with this method is that as the number of types of features such as products, attributes, etc. increases, the NMF error increases with a small number of factors. Also, if the number of factors is high, k-means clustering may not be accurate. This study proposes a method to improve the conventional method by classifying products into genres, so that shop characteristics are more accurately represented. Genre classification reduces the dimension of the input matrix and solves the aforementioned problem. Also, this method can be used to identify the similarity of each store, since order trends are simplified by genre segmentation. In experiments, we have confirmed the effectiveness of the proposed method by comparing it with the conventional method using real restaurant POS data.
In the rapidly changing world of smart technology, searching for documents has become more challenging due to the rise of advanced language models. These models sometimes face difficulties, like providing inaccurate information, commonly known as "hallucination." This research focuses on addressing this issue through Retrieval-Augmented Generation (RAG), a technique that guides models to give accurate responses based on real facts. To overcome scalability issues, the study explores connecting user queries with sophisticated language models such as BERT and Orca2, using an innovative query optimization process. The study unfolds in three scenarios: first, without RAG, second, without additional assistance, and finally, with extra help. Choosing the compact yet efficient Orca2 7B model demonstrates a smart use of computing resources. The empirical results indicate a significant improvement in the initial language model's performance under RAG, particularly when assisted with prompts augmenters. Consistency in document retrieval across different encodings highlights the effectiveness of using language model-generated queries. The introduction of UMAP for BERT further simplifies document retrieval while maintaining strong results.
Lexicon-based approaches to Document Classification are widely used, but the manual construction of lexicons can be time-consuming and resource-intensive. In this paper, we propose methods for automating the generation of lexicons later used for Document Classification. We explored diverse methods for generating lexicons, including semantic matches, frequency-based approaches, machine learning algorithms, and large language model techniques. We, later, used these lexicons to classify documents based on their content. By comparing our different lexicons results on a same task, based on criteria such as scalability and the F1 score, we determine optimized use-case for those methods. We show that our automated approaches are effective and efficient, producing accurate classifications with minimal human intervention. Some approaches have the potential to streamline the document classification process, reducing the time and resources required for manual lexicon generation, it also gives optimized use-case for the different methods. Thereafter, we discussed the obtained results.
Data related to various human movements are available due to the recent development of IoT devices and communication networks, and research on the use of historical data of human movements is being actively performed. In this study, we use log data on employee movement obtained from an entry/exit management system, which has been implemented in many businesses. We propose a method for creating a social network based on employee entry/exit history and observing employee performance. We focus on meeting room-related data from entry/exit data and evaluate employees using meeting network centrality indexes. According to the experimental results, the meeting network's closeness centrality best represents job position, while the meeting network's betweenness centrality best represents employee performance.
Smartphones and smartwatches are widespread. Devices have with sensors. Research activities to obtain data from sensors to recognize sports motion and find modifications are on the rise. These studies make it possible to provide movement feedback without a coach. However, a few studies have focused on the correctness of the movement, or which parts of the movement are correct. Herein, we propose a method to evaluate the movements of a martial arts demonstration based on sensor data obtained from a smartphone. In the demonstration, several movements, such as punching and kicking, are performed. The coach considers each movement and evaluates the overall movement comprehensively. Regarding the coach evaluation, we evaluate sensor data by dividing the data and using machine learning. This method was applied to the acceleration data of taekwondo demonstrations. It was shown to reproduce the instructor's score with high accuracy.
A social dilemma in game theory models the situation of agents who do not share a same objective in a multi-agent environment. However, since it ignores various aspects of real social dilemmas, a temporally extended sequential social dilemma (SSD) was proposed. SSD is an extension of Markov game and supposes two sets of policies, i.e., cooperative and defective. It was originally examined in two 2-dimensional grid games, but they were too complicated because we needed simulations to know what policies are in the above two sets. Therefore, we propose a new game that is SSD but has (trivial) cooperative and defective policies. It is a 1-dimensional grid game with two players. We first derive conditions of payoffs the players receive satisfying the constraints of social dilemmas and Markov games. Furthermore, we consider other payoffs that intentionally ignore the Markov property due to the nature of sequential decisions. We see the results of experiments in this game with two tabular Q-learning agents and discuss the learned policies in the game.
In recent years, communication has been conducted using various chatting tools. As the first step in analyzing communication itself, we propose a method for automatically constructing communication structures in group chats, particularly on LINE, which is widely used in Japan and other countries. First, the posts in the chat history are combined into units that represent the communication flow. Subsequently, the posts are classified into "conversation posts" and "reaction posts." A graph is constructed based on the content and interval of the posts, in which the posts are the nodes and appropriate connections are made between them. Considering a work-related group chat, we confirmed that the structure of the graph automatically constructed using the proposed method is similar to that of the manually constructed graph.
Information overload has become a problem because of the rapid development of the internet, and recommendation systems have attracted attention and research as a solution. It is difficult to use explicit feedback data such as comments and reviews, which are used in other fields because of the unique characteristics of the game field. In this study, we focused on implicit feedback data and proposed a new game recommendation system based on the relationship between game playing time, friendships, and game selection. Furthermore, we confirmed the effectiveness of the proposed method using experiments.
Binary classification and anomaly detection face the problem of class imbalance in data sets. The contribution of this paper is to provide an ensemble model that improves image binary classification by reducing the class imbalance between the minority and majority classes in a data set. The ensemble model is a classifier of real images, synthetic images, and metadata associated with the real images. First, we apply a generative model to synthesize images of the minority class from the real image data set. Secondly, we train the ensemble model jointly with synthesized images of the minority class, real images, and metadata. Finally, we evaluate the model performance using a sensitivity metric to observe the difference in classification resulting from the adjustment of class imbalance. Improving the imbalance of the minority class by adding half the size of the majority class we observe an improvement in the classifier’s sensitivity by 12% and 24% for the benchmark pre-trained models of RESNET50 and DENSENet121 respectively.
In recent years, in the field of artificial life, research has been actively conducted to model the behavior of living organisms on a computer and elucidate the evolutionary mechanism. It is known that real Aomon damselflies perform long-term mating before the start of spawning activity to restrain females in order to prevent sperm from being scraped from other males. It has also been confirmed that mating time is longer in an environment with many males. In this study, we try to create an agent-based model that strongly reflects the actual ecology in order to observe the evolution of male mating time. Then, the created model is used to investigate changes in mating time in an environment with many males.
People generally perform various activities, such as walking and running. They perform these activities with different motions. For example, walking can be performed with or without swinging shoulders, as well as staggering and swinging arms. We assume that such differences occur based on physical and mental characteristics of humans. To analyze relations between the motions and the characteristics/conditions, it is useful to group humans according to these differences. In a previous work, we proposed a method that successfully grouped humans by analyzing accelerometer data of their bodies in a specific activity with fixed timing and duration. In this study, we tackle with a problem of grouping human in generic, variable-length activities, such as walking and running. We propose a method that detects same motions from the accelerometer data with sliding windows and merges continuous same motions into a motion. The method is robust regarding the difference in timing and duration of the motion. In our conducted experiments, the proposed method classified humans into groups appropriately, the groups which are acquired by the previous method with the same data but without assuming fixed timing and fixed duration, which are assumed in the previous method. The proposed method is robust against temporally noised data generated from the data.
In recent years, the widespread use of IC cards and the development of sensor devices have made it possible to collect and store a wide range of data. Research has been conducted on analyzing human behavior using this data. Non-negative Multiple Matrix Factorization (NMMF) is one of the matrix factorization-based methods to extract frequent patterns from multiple data. Kojima et al. proposed a method for analyzing the relationship between behavioral characteristics and personal attributes using NMMF and decision tree learning on the movement log data of workers. However, there are several problems with this method of interpreting behavior patterns. Therefore, we propose a method that facilitates the interpretation of the behavioral patterns of clusters by creating and visualizing a matrix of extracted features of the patterns.
Koichi Moriyama合作论文数Sony Corporation33