
A major concern in smart power grids is when malicious or manipulated data is injected into measurement data due to malicious activities. Several approaches have been investigated to counter such false data injection attacks (FDIAs). However, such data-driven detectors present two major limitations. First, they neglect capturing the grid’s spatial characteristics. Second, they offer limited attack identification to familiar types of FDIAs since they are present within the model’s train sets. To conquer such limitations, we propose the use of an artificial intelligence-based graph autoencoder (GAE) for FDIAs detection. Our proposed detector offers three main advantages compared to existing detectors. First, it employs the operation of graph convolution to apprehend the grid’s spatial characteristics. Second, it offers an unsupervised autoencoder-based anomaly detection that requires only benign samples under normal operation for training. Third, it outperforms existing detectors by 16–47
Group imbalance data often occur in high dimensional practical classification problems where the number of attributes exceeds the number of instances. In this case, the researcher is faced with dual problems such as (i) biasedness towards the majority group over the minority group and (ii) dimensionality or singularity of the covariance matrix. In such situations, the classical classifiers dependent on the sample mean and covariance matrix are impracticable for classifications. This study focused on the effects of group imbalance data on $$n>p$$ classification problems. First, we develop a procedure that could transform the minority group into the majority group before the classifiers are applied. This study aims to determine whether there are observable effects on the classifiers’ performance for imbalanced and balanced data sets. We also investigated whether the over and under-sampling influence the computational time of the classifiers. The results revealed that the Fisher linear classification method (FLCM) performed comparably for imbalanced and balanced data and outperformed the nearest mean classifier (NMC) and the independent classification rule (ICR). The study demonstrated that the MVCT effects on the classifiers are data-dependent. Therefore, the investigation showed that sample size balancing irrespective of the data dimension does not have a strong impact on the classifier’s performance. This analysis concludes that for the $$n>p$$ classification problem, the FLCM classifier has comparable performance on the imbalanced and balanced data, similar results were observed for the NMC and ICR classifiers. The over and under-sampling of the data set has insignificant effects on the computational time of the classifiers.
Numerous studies show the ability of fetuses for affective evaluation and sensitivity to the sounds and rhythms of other human presence. This shows an appearance of fetuses’ perception in intentional engagement with the environment. It means that fetuses are able to select the relevant stimulus from the noisy environment with a cacophony of other stimuli: chemical interactions, pressure changes, and electromagnetic fields. This ability can appear in ecological learning only. The theoretical study observes the literature to understand what environmental features of the mother-fetus communication model enable a fetus to interact with the mother in ecological training. The objective is to design Human-Machine Systems and, specifically, computer-aided Medical Diagnosis systems based on the mother-fetus communication model. The article proposes the physiological mechanism of shared intentionality that relies on the mother's heart pulsed electromagnetic field (PEMF) impact on the adenosine receptors in both organisms. The study creates the concept design for future research to provide evidence of the mother-fetus communication model and establish human-computer connectivity.
This study investigates the relationship between news sentiment and the stock market's return. The sentiment was automatically analyzed using four methods, including lexicon-based and deep learning-based approaches, at three levels of granularity, i.e., sentence, paragraph, and full text. The sentiment was combined with features from the calendar year, lagged returns, and news publishers, which were fed into the XGBoost algorithm trained to classify the direction of market return for the following business day. The performance was maximized using Bayesian hyperparameter optimization and evaluated using nested cross-validation. The proof of concept was demonstrated using ten companies in the Dow Jones Index, which were grouped into five sectors. The findings indicate an asymmetric power of sentiment measures in different sectors, with the petroleum industry being the most responsive to the sentiment expressed in the news. The study highlights the significance of targeted sentiment measures in making informed decisions about the market direction, particularly for the petroleum industry.
The current technological advancements revolutionizing the concept of Urban Air Mobility (UAM), has a concurrent need to quantify the operational safety of these vehicles in terms of their associated risk. Providing safety certification of flight operations of UAM vehicles is critical as the concept relies on battery powered electrically Vertical Takeoff and Landing (eVTOL) vehicles, to operate in the current Air traffic control. In this paper, a data-driven method for UAM vehicle energy consumption prediction and risk quantification with conditional value-at-risk based on energy consumption distribution is presented. Significant factors affecting energy consumption, such as density altitude, aircraft design, airspeed, and collision avoidance algorithms, are considered in the data-driven based energy consumption prediction of multiple eVTOL flights. Additionally, a risk metric was deployed to evaluate the risk associated with worst case energy dissipating flights. Our result shows that the proposed approach provides a generalized method to quantify operational safety of UAM network over a given region.
Eye blinking has been studied extensively due to its wide range of potential applications. However, one under-researched field is the use of the wider lacrimal area for detection. This paper proposes a new eye blinking detection method using a novel lacrimal aspect ratio (LAR) strategy that utilises eyebrow movement and eyes. The proposed algorithm estimates facial landmarks using an automatic facial landmark detector to extract a single scalar quantity by using LAR and characterizing eye opening and closing, and to detect both partial and full blinking in each frame using a LAR threshold. We set three threshold values, –2.4 and –2.6, and –2.9, to detect blinks by each frame. Experimental results show that our approach successfully detects eye blinks and can outperform other state-of-the-art works. The utilization of LAR in detecting blinks and partial blinks demonstrates its potential to offer a novel and informative metric for researchers. This approach also opens up possibilities for further eye-related investigations, including the recognition of emotions. With its low dimensionality and easily understandable time domain features, LAR provides an effective pathway towards achieving these goals.
To curb energy theft and identify fraudsters and other non-technical losses, power distribution companies have used Machine Learning algorithms and a large amount of data with high granularity from electricity consumption units. Those data are collected from consumers through Advanced Metering Infrastructure (AMI) like Smart Meters (SM) being remotely collected in real time and several samples per day. In emerging countries like Brazil, most energy meters are technologically limited or electromechanical. Those devices can measure only aggregated values of energy consumption in a monthly basis. This work proposes HyMO-RF, one strategy of using a multi-objective search algorithm to improve the performance of a machine-learning model in the detection energy theft in a scenario of limited resources (AMI and SM). The proposed approach in this paper (HyMO-RF) is based on hyperparameter tuning and the main contribution is to associate multi-objective algorithms to determine an optimized combination of hyperparameters in order to maximize the classification model’s performance. A real life corporate dataset was provided by a Brazilian power distribution company CPFL ENERGIA™. The data was properly anonymized and securely stored. We used NSGAII multi-objective algorithm for hyperparameters tuning, improving the Random Forest (RF) classifier’s performance. Results achieved in terms of Precision and F1-score metrics were 0.83 and 0.73 respectively. An additional field study showed that the solution proposed already impacted the operation of the fraud detection team in a positive way by having an accuracy of 74
Due to the increasing volume of user-generated content on the web, the vast majority of businesses and organizations have focused their interest on sentiment analysis in order to gain insights and information about their customers. Sentiment analysis is a Natural Language Processing task that aims to extract information about the human emotional state. Specifically, sentiment analysis can be achieved on three different levels, namely at the document level, sentence level or the aspect/feature level. Since document and sentence levels can be too generic for an opinion estimation given specific attributes of a product or service, aspect-based sentiment analysis became the norm regarding the exploitation of user generated data. However, most human languages, with the exception of the English language, are considered low-resource languages due to the restricted resources available, leading to challenges in automating information extraction tasks. Accordingly, in this work, we propose a methodology for automatic aspect extraction and sentiment classification on Greek texts that can potentially be generalized to other low-resource languages. For the purpose of this study, a new dataset was created consisting of social media posts explicitly written in the Greek language from Twitter, Facebook and YouTube. We further propose Transformer-based Deep Learning architectures that are able to automatically extract the key aspects from texts and then classify them according to the author’s intent into three pre-defined classification categories. The results of the proposed methodology achieved relatively high F1-macro scores on all the classes denoting the importance of the proposed methodology on aspect extraction and sentiment classification on low-resource languages.
Water is more than just a necessity to sustain life on the planet by quenching the thirst of humans, animals and plants, There are many reasons why we may face a worse global water crisis in the future than we are currently experiencing. Among the most important of these reasons is the loss of large quantities of fresh water during the irrigation process. In this paper, we present a new irrigation technique that focuses on studying the stages of plant development and estimating the actual amount of water needed at each stage, in order to minimize Over-watering and Under-watering of the plant during its life stages. We use a high amount of data previously gathered through a Wireless Sensor Network (WSN) spread in different places in the agricultural field, then we use k-Nearest Neighbors (KNN) and Weighted-k Nearest Neighbors (W-KNN) to train the Machine Learning model. However, in most existing methods of irrigation the estimated amount of water directed to the plant is constant during all stages. Our proposed solution is able to overcome this disadvantage by introducing the development stages of the plant to the learning model. The results obtained through W-KNN algorithm outperform manual irrigation and automated irrigation without stages.
Augmented reality and virtual reality (AR/VR) systems contain several different sensors including image sensors for gesture recognition, head pose tracking and pupil/eye tracking. The data of all these sensors must be processed by a host processor in real-time. For future AR/VR systems, new sensing technologies are required to fulfill the demands in power consumption and performance. Currently pupil detection is performed with images on resolutions around 300 × 300 pixels and above. Therefore, deep neural networks (DNN) need host platforms, which are capable to compute the DNNs with such input resolutions to process them in real-time. In this work, the image resolution for pupil detection is optimized to a resolution of 100 × 100 pixels. A tiny pupil detection neural network is introduced, which can be processed with the ARM Cortex-M55 and the Embedded Machine Learning (ML) processor Arm Ethos-U55 with a performance of 189 frames per second (FPS) with high detection rates. This allows to reduce the power consumption of the communication between image sensor and host for future AR/VR devices.
In this paper a hybrid model is presented for generating novel stories using (a) a traditional symbolic AI cognitive-appraisal model of emotions embodied in the Affective Reasoner (AR), and (b) the large-language-model-based (LLM) system embodied in ChatGPT. The novel emotion and narrative structure is generated first by AR techniques—giving strong, symbolic computable structure to the intermediate narratives—and then fed in series to ChatGPT to add complementary world knowledge and elegant language structure. The resulting stories are polished and cohesive, but the basic structural elements remain under computational control. Explanations about content can be generated, based on the emotion content, and also on the appraisal-based dispositions, expressive temperaments, reasoning about the fortunes of others, relationships and moods of the characters in the stories. Background emotion theory is reviewed, relevant to the morphing of narratives, composed of 28 emotion categories, 24 emotion intensity variables, and ~400 channels for emotion expression, which has been implemented in the AR. A series of hybrid-generated stories are presented illustrating how the emotion makeup of characters, their emotions, their actions and their narrative perspectives remain not only consistent but are largely enhanced after treatment by ChatGPT. Actual examples of generated stories covering a wide range of complex emotion scenarios are given.
The worldwide energy-crisis poses a critical risk to the energy-intensive process industry. Rising costs for gas lead to increased usage of electrical power (e.g. for heating) that network operators are not prepared for. Weather-dependent energy-sources (e.g. windparks, solar panels) lead to additional fluctuations within the power grid. In worst case a simultaneous and prolonged loss of gas supply and electricity will lead to network bottlenecks, or complete network shutdowns—blackouts. For manufacturers, power outages thereby lead to severe consequences (i.e. waste, broken machines, additional costs), with only limited options to prevent them. Within this paper we highlight the implementation of POWOP, a weather-based service for POWer Outage Prediction that increases the resilience within the German process industry (e.g. paper, glass or chemical production). By using a predictive analytics forecasting model and a knowledge graph consisting of semantically enhanced Scenario Patterns, we are able to predict regional power outages for the next 7 days and to provide action recommendations for potential actors. Our publicly available web-application was evaluated for 15 locations of paper manufacturers in the German region Bavaria and will be demonstrated within a screencast.
Fields as diverse as art, photography, writing, and design are now confronting the consequences of easily available—and easy to use—generative systems. Text and image products that used to be understood as direct expressions of the minds of a human creator may now be generated synthetically through software and neural networks. The emergence of such 'smart' software poses novel ethical challenges for the evaluation of intellectual products like images and text. There has been a predictable backlash against incursions of automation and generative systems into creative practices that have evolved over centuries. The author surveys critical perspectives on these software systems as they relate to creative practices, and cites examples of how tools and services made possible by neural networks have caused controversy. Further, based on history and trends in mainstream digital culture, the author concludes that popular notions of authorship and creativity will continue to evolve as machine learning and artificial intelligence become increasingly entwined in production tools.
Decision Support System frameworks have great importance in the context of Industry 4.0 to prevent production bottlenecks, machine malfunction and to increase the reliability of the industrial processes environment. With the development of digitalization, Decision Support Systems (DSS) alongside Cyber Physical solutions, Internet of Things (IoT) devices and Big Data approaches constitute the main core of an industrially oriented smart manufacturing application. However, a considerable amount of industries lack the technological infrastructure in order to effectively utilize the vast amount of data, collected from various sensors and heterogenous sources scattered at various points of the production process on a daily basis. The scope of this paper is to present a conceptual framework of a DSS in a Canning Industry in order to utilize high-volume data collected from embedded sensors in the production process in order to detect and eliminate bottlenecks. The first part of the solution is dedicated to the integration of Programmable Logical Controllers (PLC) and the KEP Open Platform Communications (OPC) Server for the data acquisition and communication with a MySQL relational database, while the second part is on the data manipulation and the presentation of data analysis strategy for the decision making. Three machine learning models namely Random Forest, Naïve Bayes and SVM were tested for the prediction of total production losses with Random Forest outperforming the rest with Accuracy of 90
In this paper, we propose a new cancellation intervention system to minimize possible revenue loss of business entity in tourism sector from last-minute booking cancellation. The proposed system automatically sends e-reminders to travelers who are most likely to cancel their bookings using calibrated prediction models on the subsets of bookings with different lead times. In particular, cost-sensitive learning methods with varying class weights in machine learning community are adopted to overcome hurdles from imbalanced class distributions. Finally, this study introduces cumulative gain charts to provide general guidelines on how to maximize the expected benefits from the proposed system.
This poster documents the basic points of a research project that investigates the role of the user of AI systems in the production of meaning and content. Whether the user is a student, a researcher or an artist, the output of the human-machine cooperation should be original. As a generative assembly of textual or visual data, the outgoing content looks like the rephrasing of things that have already been said. But how AI generated texts and images surprise us? How is their production been triggered? Beginning form the assumption that the user is a compositor of thoughts and connotations, which, in cooperation with the AI system could lead to the formation of an out of the box thinking, we notice that her practice has similarities to the practices of the gardener, of the interviewer and of the interrogator. But what do these similarities concern? What do a gardener, an interviewer and an interrogator have in common with the user of AI? We'll attempt to answer these questions with the help AI itself, placing ourselves in the position of the user in question, thus creating a simulation. This is what we like to call a tautological method, since the compositing procedure replicates itself, permitting a first person observation. Whether this approach will be fruitful of not, is something that we are going to discover at the end of this poster.
Datasets that include alignments between natural language and Knowledge Graphs are fundamental to a wide variety of Natural Language Processing and Generation tasks. Current state-of-the-art aligned datasets, though, are significantly impacted by reduced size and scarcity of covered domains, and their quality is difficult to evaluate. To compensate for these issues, we introduce SEALIon, a tool for extracting RDF triples from natural language textual corpora based on a human-in-the-loop approach. We present our first results of SEALIon's approach, paving the way for further researches in the field of human-in-the-loop triple extraction.
Bitcoin prices have been predicted using Twitter sentiments, with results showing relatively low prediction accuracy. Additional external data sources, such as Google Trends, have been used to improve prediction accuracy. However, to the best of our knowledge, no analytical approach has been used to explain why Twitter sentiment is not a good predictor and why additional external data improved predictive model accuracy. Consequently, this paper uses cluster analysis and Shapley Additive Explanations (SHAP) to analyse feature importance and impact on the prediction outcome of the eXtreme Gradient Boosting (XGBoost) model. A combination of Twitter sentiments and user interaction behaviour, such as likes, retweets, and replies, are used as input variables for the XGBoost model and Bitcoin closing prices as the target variable. Our findings indicate that the sentiment score is insufficient because the majority of Bitcoin-related tweets come from Bitcoin enthusiasts whose opinions are unaffected by market fluctuations, and the improved prediction accuracy observed when external data are used in addition to the sentiment score is significant only during price volatility and can be attributed to an increase in the total number of interactions from new sets of users and not the cumulative user behaviour.
Named Entity Recognition (NER) serves as the foundation for several natural language applications like question answering, chatbots and intent classification. Identification of entity boundaries and its categorization into entity types poses a significant challenge in domain-dependent and low-resource settings, with limited training data availability. To this end, we propose AtEnA, a novel NER framework utilizing entity class attributes from external knowledge source for few-shot learning. We use a two-stage fine-tuning process, wherein a language model is initially trained to “attend” to the different entity class attributes along with the textual context, and is then fine-tuned for the downstream application data with few annotated training examples. Experiments on benchmark NER datasets depict AtEnA to perform around 10 F1 score points better than the existing NER methodologies, specifically for few-shot limited training scenarios.
The application of Explainable AI (XAI) in time series forecasting has gradually attracted attention, given the widespread implementation of machine learning and deep learning. ShapTime - A general XAI approach based on Shapley Value specially developed for explainable time series forecasting, which can explore more plentiful information in the temporal dimension, instead of only roughly applying traditional XAI approaches to time series forecasting as in previous works. Its novel components include: (1) It provides the relatively stable explanation in the temporal dimension, that is, the explanation result can reflect the importance of time itself, which is more suitable for time series forecasting than traditional XAI approaches; (2) It builds the practical application scenario of XAI - improving forecasting performance guided by explanation results. This is distinctly different from previous works, which only present the results of XAI as the demonstration of innovation. Eventually, in five real-world datasets, ShapTime's average performance improvements for Boosting, RNN-based and Bi-RNN-based reached 18, 20 and 35%, respectively.