The “HPI Future SOC Lab” is a cooperation of the Hasso Plattner Institute (HPI) and industry partners. Its mission is to enable and promote exchange and interaction between the research community and the industry partners. The HPI Future SOC Lab provides researchers with free of charge access to a complete infrastructure of state of the art hard and software. This infrastructure includes components, which might be too expensive for an ordinary research environment, such as servers with up to 64 cores and 2 TB main memory. The offerings address researchers particularly from but not limited to the areas of computer science and business information systems. Main areas of research include cloud computing, parallelization, and In-Memory technologies. This technical report presents results of research projects executed in 2017. Selected projects have presented their results on April 25th and November 15th 2017 at the Future SOC Lab Day events.
PurposeThis paper aims to describe requirements for a model that can assist in identity deception detection (IDD) on social media platforms (SMPs). The model that was discovered demonstrates the usefulness of the requirements. The aim of the model is to identify humans lying about their identity on SMPs.Design/methodology/approachThe requirements of a model for IDD will be determined through a literature study combined with a study that identifies currently available identity related metadata on SMPs. This metadata refers to the attributes that describe a user account on an SMP. The aim is to restrict IDD to be only based on these types of attributes, as opposed to or combined with the contents of a single or multiple communications.FindingsData science experiments were conducted and in particular supervised machine learning models were discovered that indeed detects identity deception on SMPs with an area under the receiver operator characteristics curve (ROC-AUC) of 75.5 per cent.Originality/valueSMPs allow any user to easily communicate with their friends or the general public at large. People can now be targeted at great scale, most often for malicious purposes. The reality is that many of these cyber-attacks involve some form of identity deception, where the attackers lie about who they are. Much focus to date has been on the identification of non-human deceptive accounts. This paper focuses on deceptive human accounts that target vulnerable individuals on SMPs.
Social media platforms allow billions of individuals to share their thoughts, likes and dislikes in real-time, without any censorship. This freedom, however, comes at a cyber-security risk. Cyber threats are more difficult to detect in a cyber world where anonymity and false identities are ever-present. The speed at which these deceptive identities evolve calls for solutions to detect identity deception. Cyber-security threats caused by humans on social media platforms are widespread and warrant attention. This research posits a solution towards the intelligent detection of deceptive identities contrived by human individuals on social media platforms (SMPs). Firstly, this research evaluates machine learning models by using attributes such as the “profile image” found on SMPs. To improve on the results delivered by these models, past research findings from the field of psychology, such as that humans lie about their gender, are used. Newly engineered features such as “gender-derived-from-the-profile-image” are evaluated to grasp whether these features detect deception with greater accuracy. Furthermore, research results from detecting non-human (also known as bot) accounts are also leveraged to improve on the initial results. These machine learning results are lastly applied to a proposed model for the intelligent detection and interpretation of identity deception on SMPs. This paper shows that the cyber-security threat of identity deception can potentially be minimized, should the vulnerability in the current way of setting up user accounts on SMPs be re-engineered in the future.
Social Media Platforms (SMPs) allow any person to easily communicate with their friends or the general public at large. People can now be targeted at great scale, most often for malicious purposes. The mere fact that more people are using SMPs exposes more people to various forms of cyber threats such as cyber-bullying. The problem is that many of these cyber-attacks involve some form of identity deception, where the attackers lie about who they are. The solution proposed in this paper is to work towards developing a model for Identity Deception Detection (IDD) on SMPs by identifying and using metadata that is freely available on SMPs. This metadata includes attributes that describes a user account on an SMP. The aim is to use only these attributes, as opposed to the contents of a communication, for determining if people are lying about their identities. By discarding contents, an identity deception detection model can be developed with lower overhead. A prototype is discussed that runs an experiment using the metadata (attributes) that defines the identity of a user on an SMP. The results show promise for further research in developing solutions for assisting with the automatic detection of identity deception.
There are a growing number of people who hold accounts on social media platforms (SMPs) but hide their identity for malicious purposes. Unfortunately, very little research has been done to date to detect fake identities created by humans, especially so on SMPs. In contrast, many examples exist of cases where fake accounts created by bots or computers have been detected successfully using machine learning models. In the case of bots these machine learning models were dependent on employing engineered features, such as the "friend-to-followers ratio.'' These features were engineered from attributes, such as "friend-count'' and "follower-count,'' which are directly available in the account profiles on SMPs. The research discussed in this paper applies these same engineered features to a set of fake human accounts in the hope of advancing the successful detection of fake identities created by humans on SMPs.
The bulk of currently available research in identity deception focuses on understanding the psychological motive behind persons lying about their identity. However, apart from understanding the psychological aspects of such a mindset, it is also important to consider identity deception in the context of the technologically integrated society in which we live today. With the proliferation of social media, it has become the norm for many people to present a false identity for various purposes, whether for anonymity or for something more harmful like committing paedophilia. Social media platforms (SMPs) are known to deal with massive volumes of big data. Big data characteristics such as volume, velocity and variety make it not only easier for people to deceive others about their identity, but also harder to prevent or detect identity deception. This paper describes the challenges of identity deception detection on SMPs. It also presents attributes that can play a role in identity deception detection, as well as the results of an experiment to develop a so-called Identity Deception Indicator (IDI). It is believed that such an IDI can assist law enforcement with the early detection of potentially harmful behaviour on SMPs.
Das Future SOC Lab am HPI ist eine Kooperation des Hasso-Plattner-Instituts mit verschiedenen Industriepartnern. Seine Aufgabe ist die Ermoglichung und Forderung des Austausches zwischen Forschungsgemeinschaft und Industrie. Am Lab wird interessierten Wissenschaftlern eine Infrastruktur von neuester Hard- und Software kostenfrei fur Forschungszwecke zur Verfugung gestellt. Dazu zahlen teilweise noch nicht am Markt verfugbare Technologien, die im normalen Hochschulbereich in der Regel nicht zu finanzieren waren, bspw. Server mit bis zu 64 Cores und 2 TB Hauptspeicher. Diese Angebote richten sich insbesondere an Wissenschaftler in den Gebieten Informatik und Wirtschaftsinformatik. Einige der Schwerpunkte sind Cloud Computing, Parallelisierung und In-Memory Technologien. In diesem Technischen Bericht werden die Ergebnisse der Forschungsprojekte des Jahres 2015 vorgestellt. Ausgewahlte Projekte stellten ihre Ergebnisse am 15. April 2015 und 4. November 2015 im Rahmen der Future SOC Lab Tag Veranstaltungen vor.
Identity Deception Detection is a problem on social media platforms today. Not only is there challenges towards determining the authenticity of people, but also with analyzing the data that forms part of the communications. These data are of heterogeneous type and include photos, videos and sound. Furthermore, most social media platforms are operating in an uncontrolled environment. Any person can contribute content and take part. Even though age restrictions do exist there are no enforcement of these laws and honesty of the public is expected. This is dangerous for minors specifically as they are either unaware of the dangers or not mature enough to be responsible for their actions online. Online predators are aware of this fact and targeting this group specifically. This paper presents work-in-progress towards developing an intelligent Identity Deception Indicator (IDI). It is envisaged that this work could eventually assists authorities in doing large-scale observation on publicly available social media platforms, such as Twitter. Of particular interest are those personas whose behavior and online content does not fit with the age group they are conversing with.