Detection of cyber grooming is important, both in a forensic and an early detection setting. Many communication apps now provide end-to-end encryption to safeguard the privacy of the communication. This means that for grooming detection the messages are no longer available. Under the assumption that we are able to still use the messages of a child can be extracted from its typing on the keyboard, we have investigated the impact on performance of cyber grooming detection when only half of the conversation is available. We found that we can get comparable performance in both a forensic setting and an early detection setting and in some cases the performance when only using the victim’s messages for training and testing improved when compared to using the messages of both chatters for training and testing.
As more industries employ robots to perform critical tasks, the need to secure such robots are increasing. Mobile robots are more vulnerable to being attacked, as these are not always deployed in well-controlled environments. Therefore, security of mobile robots is essential, where cyber- and physical-attacks can lead to catastrophic events like physical injury, even loss of life. In some cases, such mobile robots include limited to no security controls. Preventing such cyber-attacks is not always possible. However, timely detection of such attacks or attempted attacks might lead to the deployment of appropriate response actions, which limit the negative consequences of the attack and potentially block corresponding attack-vector. When considering intrusion detection in mobile robots, it is necessary to monitor both the cyber and physical domain, as by their nature, attacks conducted in the cyber realm can lead to severe damage in the physical realm. However, there is limited research done on how such cyber and physical attacks against mobile robots can be detected. Therefore, in this paper, we developed a system for intrusion detection in mobile robots, using a Machine Learning-based approach. This developed system can detect cyber- and physical-attacks, even when the robot is deployed in a previously unknown environment. Our proposed system shows promising results when evaluated utilizing datasets that we collected from Spot robot. To assess the performance of this system, we performed two physical- and two cyber-attacks against the robot, which were identified as a part of the review on threat landscape for mobile robots. This system can be applicable to all mobile robots, to detect attacks in both the cyber- and physical-domain, as the data used for intrusion detection in the context of this study should be available in other mobile robots.
This book is on biometric authentication and proposes to use keystroke dynamics to avoid password-based authentication problems such as shared or stolen. The difficulties concerning password-based are that most users option for simple passwords. They prefer using similar passwords spanning distinct applications (Vance, 2010). Keystroke dynamics is known to overcome these circumstances. Keystroke dynamics measures the rhythms that a person exhibits while typing on a keyboard. In this sense, keystroke dynamics is a behavioural biometric modality, as well as signature dynamics, gait and voice (Klevans and Rodman, 1997; Monrose and Rubin, 2000; Impedovo and Pirlo, 2007; Moustakas et al., 2010). Among the advantages of keystroke dynamics in comparison to other modalities, it can be mentioned here that it is a low-cost modality: indeed, neither extra sensor nor device is required (Giot et al., 2011; Bours, 2012). The counterpart to this low cost and ease of use is the worst performances compared to those obtained with morphological biometric modalities such as fingerprint, face and iris (Wildes, 1997; Maio and Jain, 2009). The rather worse performances of keystroke dynamics (in comparison to other modalities) can be explained by the large intra-class variability of the users’ behaviour. Indeed, the way of typing continuously evolves when time elapses. One way to handle this variability is to consider additional information in the decision process. Future perspectives of these security procedures could be a new paradigm in cybersecurity applications.
eHealth systems require usable but more robust authentication mechanisms to balance security and usability. Continuous authentication is a security mechanism that passively conducts user authentication throughout the session. Continuous authentication may best fit healthcare systems as it enhances security and improves usability by seamlessly authenticating users. It may face limitations when only one modality is supported, such as keystroke dynamics, gait dynamics, touch dynamics, etc. These modalities collect and utilize user-sensitive data containing information about user behavioral and contextual activities, and other user-sensitive attributes, e.g., user gender, age, etc., may also be derived from such data, which causes privacy concerns. Continuous authentication using multiple modalities may overcome the limitations of a single modality at the cost of compromising user privacy. The more modalities we employ, the more privacy we compromise. In this paper, we propose a privacy-preserving protocol that supports continuous authentication using multiple modalities. Our proposed protocol protects 1) user-sensitive attributes and 2) the privacy of the type of modality (such as user activities). The biometric performance of the proposed protocol is determined in the following ways: a) individually, on two public datasets, a keystroke dynamics dataset, and a swipe gesture dataset, and b) multimodal, by combining swipe gesture and keystroke data. For multimodal, instead of computing cosine similarity for each action, we comput ed the extended similarity based on multiple (k) keystroke and swipe gesture actions. The experimental evaluation proves that our proposed protocol with the extended technique performs better than the original cosine similarity. The proposed protocol offers efficient biometric performance, low communication and computation costs, and security in the presence of a semi-honest authentication server, malicious users, and external adversaries.
This article investigates modern approaches to preventing online child grooming, emphasizing raising awareness and analyzing current trends. It focuses on identifying emerging threats and examining existing prevention methods and available data within established models. Based on these findings, a multilingual web application was developed at the Technical University of Košice to protect children from online grooming in Slovakia and beyond. The web application allows users to assess online conversations and statements to determine whether they are safe or show signs of grooming, vulgarity, blackmail, or other related behaviors. The platform is available in three languages: English, Slovak, and Norwegian. The results from testing the app, collected from unique users, are presented and analyzed in the final section.
This paper investigates the impact of using Google Translate to perform authorship attribution analysis on Norwegian texts translated into English, with the goal of detecting astroturfing activities in online public discourse. The study compares the performance of various n-gram-based attribution techniques on original Norwegian texts and their machine-translated English counterparts. Results show that the best performance was achieved without translation or pre-processing of the original Norwegian texts, reaching an accuracy of 0.84. Further research is needed to explore alternative methods and pre-processing techniques to enhance the accuracy of authorship on translated texts in this context.
The further development of human–computer interaction applications is still in great demand as users expect more natural interactions [...]
Social media platforms present significant threats against underage users targeted for predatory intents. Many early research works have applied the footprints left by online predators to investigate online grooming. While digital forensics tools provide security to online users, it also encounters some critical challenges, such as privacy issues and the lack of data for research in this field. Our literature review investigates all research papers on grooming detection in online conversations by looking at the psychological definitions and aspects of grooming. We study the psychological theories behind the grooming characteristics used by machine learning models that have led to predatory stage detection. Our survey broadly considers the authorship profiling research works used for grooming detection in online conversations, along with predatory conversation detection and predatory identification approaches. Various approaches for online grooming detection have been evaluated based on the metrics used in the grooming detection problem. We have also categorized the available datasets and used feature vectors to give readers a deep knowledge of the problem considering their constraints and open research gaps. Finally, this survey details the constraints that challenge grooming detection, unaddressed problems, and possible future solutions to improve the state-of-the-art and make the algorithms more reliable.
Detecting predatory behavior in online conversations is crucial for ensuring the safety of its users, in particular minors. This study investigates the performance of various machine learning models in identifying predatory conversations using a sliding window approach, that looks at a set number of messages in an online conversation. The models are evaluated based on their accuracy, precision, recall, and F 1 scores. The results showed that certain models performed differently with different window sizes. The study has some limitations, such as small dataset sizes and fixed sliding window sizes. This warrants further investigation, and future research should focus on evaluating the models on a bigger dataset, exploring other window sizes, and considering different hyper-parameters for the evaluation of the risk-scoring.
Depression is a serious mental health condition that affects a person’s ability to feel happy and engaged in activities. The COVID-19 pandemic has led to an increase in depression due to factors such as isolation, financial stress, and uncertainty about the future. Additionally, restrictions on travel and socializing have contributed to feelings of loneliness and isolation. In this research, we present a deep learning framework named CoDeS (COVID-caused depression symptoms) for detecting prodromes of depression in online users caused due to COVID pandemic. This framework uses a combination of CNN, LSTM, and integrated CNN-LSTM techniques, with three different feature representation methods, viz. Word2Vec, TF-IDF, and BERT. Nine experiments were conducted on individual and integrated models, and the results were evaluated based on the accuracy, precision, recall, F1-score, and Matthews correlation coefficient (MCC) performance metric. The highest accuracy value of 98.95% was recorded for the TF-IDF-based integrated CNN+LSTM model. When the same integrated model was trained using Word2Vec and BERT-based features, it still performed well with an accuracy of 97.32% and 98.51% respectively. The results demonstrate that the TF-IDF-based feature representation performed better than the Word2Vec and BERT-based feature representations for the CNN and LSTM models in identifying COVID-caused depression symptoms. The proposed approaches showcased substantial advancements over the existing ones, with significant improvements in accuracy. TF-IDF-CNN+LSTM achieved an accuracy approximately 37.28% higher, while BERT-CNN, BERT-LSTM, and BERT-CNN+LSTM achieved accuracy enhancements of approximately 29.78%, 34.44%, and 27.14% respectively. These accuracy improvements demonstrate the superior classification capabilities of the proposed approaches, leading to more precise depression analysis outcomes. In terms of F1 measure, the proposed approaches consistently demonstrated superior performance, with F1 measure values ranging from 0.965 to 0.987. BERT-CNN+LSTM achieved the highest F1 measure, highlighting its balanced precision and recall. Overall, the proposed approaches outperformed existing ones in terms of recall, precision, accuracy, and F1 measure, with improvements ranging from 27.14 to 44.85%. By incorporating advanced techniques such as TF-IDF, CNN, LSTM, and BERT, more accurate and reliable sentiment analysis outcomes can be achieved, offering the potential for enhanced applications in this field.
Cyber grooming is a compelling problem worldwide nowadays since people spend most of their time online. All of the reports strongly suggested that it becomes very urgent to tackle the online child grooming problem in order to protect children from sexual exploitation. Automatic sexual predator identification can be a promising solution to this issue since the number of online conversations is too large to be monitored manually. In this work, a two-stage is proposed with a combination of several features. The first stage is for detecting the predatory conversations while the second step aims to distinguish the predator from the victim in the predatory conversations. The features ensemble used will be combining lexical and behavioral features. The lexical features used include BoW, POS-based, topical, and emotion-based. Meanwhile, the behavioral features used for this work include the number of messages, the average number of words, the number of exclamation marks, the number of questions, sentence complexity and readability, and the number of intentions. SVM was used as a classifier due to its good ability for many text classification tasks. The experiment result shows that BoW with tf - idf term weighting provided the best performance for both PCI and VPD tasks. BoW with tf- idf term weighting obtained an F 0 .5 score of 0.9893 on PCI and 0.9798 on VPD. The features ensemble can exceed most of the individual features that form it but still cannot beat BoW.
Protecting children against predators that use the internet as a medium to find their victims is essential, where the difficulty of monitoring online messaging platforms to avoid potential threats toward underage users is alarming. Online grooming detection requires a deep knowledge of predatory behaviour where an online conversation's positive or negative connotations rely on the context where it can change rapidly in an online chat. Therefore, it is essential to use robust feature vectors for predatory conversation detection. This research paper proposes a contrastive learning framework for feature extraction in a sentence-based manner where it can assign a feature vector to the conversation with misspellings using subword information. Also, it is vital to have a high detection rate of true positives to avoid potential non-wanted consequences. We propose a configuration of RoBERTa encoders and supervised SimCSE for training the SVM model, leading to a high rate of detecting relevant samples (predatory samples). Our proposed approach gains an F0.5-score of 0.96, an F1-score of 0.96, and an accuracy of 0.99 for predatory conversation detection that benchmarks the state-of-the-art. In order to improve the performance, we also perform experiments with various fusion approaches, where the sum fusion of all configurations obtains an accuracy of 0.99, an F_1-score of 0.97, and an F0.5-score of 0.98.
In this article, we describe a method for determining a person’s real age based on voice recognition. We seek to contribute to the realm of age group classification not reliant on text by investigating voice feature analysis, making use of accessible voice audio datasets, and employing a prototype classification model. Additionally, we aim to explore the initial phase of voice-based age group classification to enhance the customization of interventions.
User authentication based on muscle tension manifested during password typing seems to be an interesting additional layer of security. It represents another way of verifying a person’s identity, for example in the context of continuous verification. In order to explore the possibilities of such authentication method, it was necessary to create a capturing software that records and stores data from EMG (electromyography) sensors, enabling a subsequent analysis of the recorded data to verify the relevance of the method. The work presented here is devoted to the design, implementation and evaluation of such a solution. The solution consists of a protocol and a software application for collecting multimodal data when typing on a keyboard. Myo armbands on both forearms are used to capture EMG and inertial data while additional modalities are collected from a keyboard and a camera. The user experience evaluation of the solution is presented, too.
Predatory conversation detection on social media can proactively prevent the netizens, including youngsters and children, from getting exploited by sexual predators. Earlier studies have majorly employed machine learning approaches such as Support Vector Machine (SVM) for detecting such conversations. Since deep learning frameworks have shown significant improvements in various text classification tasks, therefore, in this paper, we propose a deep learning-based classifier for detecting predatory conversations. Furthermore, instead of designing the system from the beginning, transfer learning has been proposed where the potential of the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model is utilized to solve the predator detection problem. BERT is mostly used to encode the textual information of a document into its context-aware mathematical representation. The inclusion of this pre-trained model solves two major problems, i.e. feature extraction and Out of Vocabulary (OOV) terms. The proposed system comprises two components: a pre-trained BERT model and a feed-forward neural network. To design the classification system with a pre-trained BERT model, two approaches (feature-based and fine-tuning) have been used. Based on these approaches two solutions are proposed, namely, BERT_frozen and BERT_tuned where the latter approach is seen performing better than the existing classifiers in terms of $$F_1$$ and $$F_{0.5}$$ -scores.
The recent experience in the use of virtual reality (VR) technology has shown that users prefer Electromyography (EMG) sensor-based controllers over hand controllers. The results presented in this paper show the potential of EMG-based controllers, in particular the Myo armband, to identify a computer system user. In the first scenario, we train various classifiers with 25 keyboard typing movements for training and test with 75. The results with a 1-dimensional convolutional neural network indicate that we are able to identify the user with an accuracy of 93% by analyzing only the EMG data from the Myo armband. When we use 75 moves for training, accuracy increases to 96.45% after cross-validation.
Contract cheating has become a profound issue in academics with the onset of the COVID-19 pandemic as digitised evaluation has become common practice. This evaluation method opens up for examining students remotely, either by online home exams or longer written assessments done away from the classroom. Contract cheating refers to a problem where the students hire a third party to complete their assignment and submit it for grading as their own. Manually dealing with contract cheating is a cumbersome task and tools for plagiarism detection are not able to detect contract cheaters as students do not use the work of other authors without consent. In this paper, a machine learning based system is designed to specifically detect the cases of contract cheating in academics. The system uses keystroke biometric behaviour where typing style is analysed to discriminate cheaters from genuine students. The experiments are conducted on two datasets where one is existing and another is designed by performing data collection in a university for recording the keystroke features. Two categories of keystroke dynamics, namely duration and latency-based features are studied for designing the various machine learning-based systems for investigating the efficient one. Furthermore, the performance of the systems are evaluated under the setting of zero false accusations in order to avoid genuine students being charged as imposters.
This study designs an online sexual predator detection system using Social Behavior Biometric (SSB) features. Social biometric focuses on extracting the pattern a user exhibits while interacting and communicating through social networks. The paper addresses the online sexual predator problem by mining the vocabulary and emotional behavior, which could assist in identifying if the user is a benign or predator. The feature-set consists of vocabulary terms that appear differently in predator and victim content. In order to strengthen the detection model, the paper also focuses on distinguishing the two classes of users based on emotions reflected in their conversation. The experiments are performed on the PAN 2012 corpus. Two datasets are created with respect to vocabulary-based and emotion-based features. The results obtained on the test set have proved that by integrating the vocabulary and emotion-based attributes, the performance of the system is significantly enhanced. While comparing, the proposed approach has outperformed top existing methods by obtaining F1, F2, and F0.5 values of 0.95, 0.94, and 0.96 respectively. Furthermore, we also recorded the best accuracy compared to state-of-the-art studies for our proposed SBB-based approach with 99.86%, 99.51%, and 99.88% for Decision Tree (DT), Support Vector Machine (SVM), and Random Forest (RF) respectively.
Automatic Gender Classification (AGC) is an essential problem due to its growing demand in commercial applications, including social media and security environments such as the airport. AGC is a well-researched topic both in the field of Computer Vision and Biometrics. In this paper, we propose the use of decision-level fusion for AGC in videos. Our approach does a decision-level fusion of labels obtained from two fine-tuned deep-networks based on a color image and optical-flow image, respectively, based on Resnet-18 architecture. We compare our proposed method with handcrafted features, which includes the concatenation of the Histogram of Optical Flow (HOF) and Histogram of Oriented Gradients (HOG). We compare it with deep-networks, which includes pre-trained & fine-tuned Resnet-18 based on a color image, and pre-trained & fine-tuned Resnet-18 based on optical flow image. Our fusion-based approach considerably outperforms both the handcrafted features, and the deep-networks previously mentioned. Another advantage of our proposed method is that it can work when the visual features are hidden. We used 98 videos from the HMDB51 action recognition dataset, specifically from the cart-wheel action with an almost 50% training, and testing split without validation set. We achieve an overall accuracy of 79.59% with Resnet-18 network architecture with the proposed method, compared to fine-tuned single-stream Resnet-18 for Color-Stream at 65.30%, and Optical-Flow at 55.10% respectively.
The abundant dissemination of misinformation regarding coronavirus disease 2019 (COVID-19) presents another unprecedented issue to the world, along with the health crisis. Online social network (OSN) platforms intensify this problem by allowing their users to easily distort and fabricate the information and disseminate it farther and rapidly. In this paper, we study the impact of misinformation associated with a religious inflection on the psychology and behavior of the OSN users. The article presents a detailed study to understand the reaction of social media users when exposed to unverified content related to the Islamic community during the COVID-19 lockdown period in India. The analysis was carried out on Twitter users where the data were collected using three scraping packages, Tweepy, Selenium, and Beautiful Soup, to cover more users affected by this misinformation. A labeled dataset is prepared where each tweet is assigned one of the four reaction polarities, namely, E (endorse), D (deny), Q (question), and N (neutral). Analysis of collected data was carried out in five phases where we investigate the engagement of E, D, Q, and N users, tone of the tweets, and the consequence upon repeated exposure of such information. The evidence demonstrates that the circulation of such content during the pandemic and lockdown phase had made people more vulnerable in perceiving the unreliable tweets as fact. It was also observed that people absorbed the negativity of the online content, which induced a feeling of hatred, anger, distress, and fear among them. People with similar mindset form online groups and express their negative attitude to other groups based on their opinions, indicating the strong signals of social unrest and public tensions in society. The paper also presents a deep learning-based stance detection model as one of the automated mechanisms for tracking the news on Twitter as being potentially false. Stance classifier aims to predict the attitude of a tweet towards a news headline and thereby assists in determining the veracity of news by monitoring the distribution of different reactions of the users towards it. The proposed model, employing deep learning (convolutional neural network(CNN)) and sentence embedding (bidirectional encoder representations from transformers(BERT)) techniques, outperforms the existing systems. The performance is evaluated on the benchmark SemEval stance dataset. Furthermore, a newly annotated dataset is prepared and released with this study to help the research of this domain.
Christophe Rosenberger合作论文数ENSICAEN - GREYC5
Einar Snekkenes合作论文数Teknologivn. 22, P.O. Box 191, 2802 Gjøvik, NORWAY4
Stephen Wolthusen合作论文数Information Security Group, Department of Mathematics
Royal Holloway, University of London3