The rising popularity of machine learning has resulted in quality data becoming increasingly valuable. However, in some cases, the data are too sparse to effectively train an algorithm or the data cannot be disclosed to unaffiliated researchers due to privacy concerns. The sparsity of data may also affect various data analyses that require a certain volume of data to be accurate. One possible solution to the aforementioned problems is data generation. However, to be a viable solution, data generation must simulate real-life data well. To this end, this paper tests whether a previously presented iterative data generation method that generates synthetic data sets based on the attribute distributions and correlations of a real-life data set can faithfully reproduce a clustered data set. The approach is shown to be ineffective for the proposed application, and consequently, a new method is introduced that might preserve the clusters present in the real-life data set. The new method is demonstrated to not only preserve the clusters within the synthetic data set, but also improve the similarity of the attribute correlations of the synthetic data set and the real-life data set.
Population census is often a popular topic in everyday conversations. After extensive research, a conclusion was made that currently no services exist that offer a convenient, easily interpretable information about the census. Our solution for this problem is CroStats – a simple web interface for population census data visualization intended for general public. The central feature of the CroStats is its intuitive and simple graphical display. It analyzes the most important categories of the population census by county such as population, age, birth rate and mortality. All this information is displayed on an interactive map of the Republic of Croatia. Extra features offered are graphs of changes by year and interesting historical and demographic facts about counties that the broader audience may have interest in.
Data warehouses are an important part of decision support systems in business. The volume of data currently being created can at times push the capabilities of relational data warehouses to their limits. A possible step forward is to use NoSQL solutions to model data warehouses, since they were made for the ever-increasing amount of data that various platforms deal with. However, simply deciding a data warehouse should be based on a NoSQL approach does not mean the problem has been solved. The flexibility of NoSQL leads to a host of new problems, such as how to perform various OLAP operations on a data warehouse that does not have a fixed schema or how and when to compute aggregate values. This paper provides an overview of various solutions that have been theorized and presented along with their advantages over relational data warehouses, as well as their drawbacks.
Nowadays, data created through the usage of different services are most commonly not available to the average researcher. Security and privacy have become a top concern, which has further restricted access to certain real-life data, especially holding true for social networks. This is why synthetic data generators have become a very important area of research, particularly synthetic social graph generators. However, even today, such generators mostly create graphs that contain just the information whether two nodes are connected. Fortunately, there is an existing conceptual solution for an expanded social graph generator that aims to generate synthetic graphs containing multiple weighted edges between nodes, thus showing various types of relationships among those nodes, all based on known real-life data characteristics. One of its proposed steps is the generation of necessary data according to provided distributions and correlations. This paper focuses on the generation of such data by adapting an existing iterative algorithm for non-normal multivariate data simulation to generate synthetic data based on the publicly available distributions and correlations of Facebook interaction parameters. It is shown that the characteristics of the generated synthetic data are similar to the known characteristics of the real-life data, proving that the chosen algorithm, along with the accompanying alterations, can be used as one of the steps within the process of generating a synthetic expanded social graph.
Social networks have long been the subject of scientific researches, frequently hindered by the unavailability of representative datasets. The advent of online social networks (OSNs), which store data about interactions between billions of people, has greatly alleviated this problem. Since user interaction on the OSNs can correspond to their real-life relationships, OSN datasets quickly became a highly sought-after resource for social network research. However, enabling open access to such data entails serious security and privacy risks, especially after the introduction of the European General Data Protection Regulation. Some researchers mitigate this problem through anonymization, while others argue for the creation of synthetic datasets. We consider synthetic datasets preferable since they circumvent the security and privacy issues. Existing synthetic dataset generators produce a social graph containing only information whether a pair of nodes are connected. However, interpersonal relationships are much more complex. Because of that, our research considers the possibility of generating a synthetic expanded social graph which, in addition to the information about the existence of a connection between a pair of users, also provides information about the types and intensities of users' interactions. As Facebook is the leading OSN today, we performed an extensive analysis of Facebook users' interaction records with the aim of getting an insight into real-life interaction patterns. In this paper, we present results of this analysis at ego-user level and offer the conceptual solution for synthetic expanded social graph generation, which uses conducted analysis results as its basis.
Online social networks (OSN) are one of the most popular forms of modern communication and among the best known is Facebook. Information about the connection between users on the OSN is often very scarce. It's only known if users are connected, while the intensity of the connection is unknown. The aim of the research described was to determine and quantify friendship intensity between OSN users based on analysis of their interaction. We built a mathematical model, which uses: supervised machine learning algorithm Random Forest, experimentally determined importance of communication parameters and coefficients for every interaction parameter based on answers of research conducted through a survey. Taking user opinion into consideration while designing a model for calculation of friendship intensity is a novel approach in opposition to previous researches from literature. Accuracy of the proposed model was verified on the example of determining a better friend in the offered pair.
In the last few decades sociologists were trying to explain human behaviour by analysing social networks, which requires access to data about interpersonal relationships. This represented a big obstacle in this research field until the emergence of online social networks (OSNs), which vastly facilitated the process of collecting such data. Nowadays, by crawling public profiles on OSNs, it is possible to build a social graph where "friends" on OSN become represented as connected nodes. OSN connection does not necessarily indicate a close real-life relationship, but using OSN interaction records may reveal real-life relationship intensities, a topic which inspired a number of recent researches. Still, published research currently lacks an extensive exploratory analysis of OSN interaction records, i.e. a comprehensive overview of users' interaction via different ways of OSN interaction. In this paper, we provide such an overview by leveraging results of conducted extensive social experiment which managed to collect records for over 3200 Facebook users interacting with over 1,400,000 of their friends. Our exploratory analysis focuses on extracting population distributions and correlation parameters for 13 interaction parameters, providing valuable insight into OSN interaction for future researches aimed at this field of study.
Online social networks (OSNs) are platforms which facilitate social interactions between their users through message exchange, photo and video sharing, status updates, etc. One of the most popular OSNs is Facebook. Connections between users on Facebook are modeled through concept of friendship. Each connection between users is binary — two users either are or aren't "friends". Information about of the actual intensity or nature of their connection is not available although in real life it can vary significantly. A majority of observed network friends are acquaintances in real-life while close friends are in the minority. The goal of this paper is to demonstrate and evaluate how user interaction statistics can be utilized for effective assessment of the nature of users' real-life relationship. Using an ensemble of popular classification algorithms, we will classify ego-user's network friends into 3 groups: close friends, friends and acquaintances. As our main contribution, we will compare the efficiency of chosen algorithms and suggest the best approach for conducting this type of analysis on similar OSN communication data.
Online social networks (OSN) are one of the most widely adapted services of the Internet infrastructure, Facebook being one of the most popular among them. Facebook models connections between its users through the concept of "friendship". However, the type and intensity of these connections between different people on Facebook vary significantly. In most cases, friends on Facebook correspond to mere acquaintances in real-life, with only a smaller subset representing actual close friends. The aim of research presented in this paper is to provide a method for estimating the intensity of Facebook friendships, i.e., to distinguish connections representing close friends from others. The study was performed by analyzing Facebook interactions between users (e.g. number of mutual likes, comments, shared photos, etc.) using supervised learning algorithms for binary classification of data. Among the chosen algorithms, the best results were gained by using random forest algorithm - accuracy of 84.73%.
Information systems of educational organizations often represent a potential well of useful information which can be discovered and interpreted by using specific methods. Exam results in particular are commonly used as a single-use measure of individual knowledge states, after which they are archived and subsequently never used again. Our approach suggests using past exam results as a rich data source for extracting knowledge about learning concepts, especially regarding their mutual relationships. To achieve this goal, we adopt our method for interactive visualization of patterns in transactional data and apply it to knowledge state matrices generated from real-life exam results and Q-matrices constructed by domain experts, providing the end user with rich, easily interpretable and visually engaging dendrogram structures.
Social learning is a natural way of acquiring knowledge and approaching various problems. In high education, students widely use social learning which is to a large extent facilitated by todays' popular social network applications. However, as of yet high education systems typically do not recognise the value of the data produced by or embedded within these social network applications. In this paper we present survey results through which we examine certain recent learning habits among students in Croatia. We predominantly investigated their usage of Facebook, today's leading social network application, specifically for the purposes of learning. Based on these findings, and taking into account common communication problems we perceive in our work, we propose a framework for University Social Network construction. We presume various usages of it by different stakeholders, propose node construction within the network compliant with educational and research processes and recommend flexible privacy levels for communication units.
Facebook, the popular online social network, is used on daily basis by over 1 billion people each day. Its users use it for message exchange, sharing photos, publishing statuses etc. Recent research shows that a possibility exists of determining the connection level (or strength of their friendship, tie strength) between users based on analyzing their interaction on Facebook. The aim of this paper is to explore, as a proof of concept, the possibility of using a model for calculating strength of friendship to compare and classify ego-user’s Facebook friends. A survey, which involved more than 2500 people and collected a significant amount of data, was conducted through a developed web application. Analysis of collected data revealed that the model can determine with a high level of accuracy the stronger connections of ego-user and classify ego-user’s friends into several groups according to the estimated strength of their friendship. Conducted research is the base for creating an enriched social graph – graph which shows all kinds of relations between people and their intensity. Results of this research have plenty of potential uses, one of which is specifically improvement of the education process, especially in the segment of e-learning.
After joining the European Union, the Republic of Croatia became obliged to consolidate its laws and regulations with the EU's. European Commission's Regulation No 260/2012, obligatory for Croatia too, has set the 31st of October 2016 as the deadline for replacement of national euro credit transfer and direct debit schemes by SEPA Credit Transfer (SCT) and SEPA Direct Debit for the EU member countries that do not use euro as their currency. From that date, use of SCT will be obligatory for the credit transfers in euro, but it has been decided that it will be obligatory for national currency kuna, too. Since SEPA is based on ISO 20022 norm, which is built upon XML language, it is not realistic to expect from business subjects, i.e. future SEPA users, the understanding of XML technologies and ISO 20022 norm necessary for creating SCT electronic payment order. For that reason, we have built a web application, which provides for users a simple web interface with forms to fill and facilitate their transit towards SCT. This paper explains SEPA and describes its implementation in Croatia, structure of SCT electronic payment initiation order format and the web application we designed for its creation.
Elektronicko poslovanje je oblik poslovanja nastao uslijed razvoja novih informacijskih i komunikacijskih tehnologija s ciljem modernizacije i olaks
In today's world, social networks are one of the most popular ways of communication. Communication and relations among people can be monitored based on interaction on social networks. The question is the exact meaning of this interaction and can the real life relationship be interpreted from interaction on social networks. In this study, we used the most popular social network Facebook with the aim of finding the correlation between the interaction among users on Facebook and friendship in real life. In this paper, we propose a model for calculating the weight of friendship among users of social networks based on their interactions on Facebook. The model takes into account the general significance of a particular form of interaction (like, comment, etc.) and the specific significance of this form of interaction for each user. Apart from creating a model for calculating the weight of friendship, in this paper, general significance ratio of each communication parameter was experimentally determined. The model was built and evaluated by searching intersection of two sets, a set of user's 10 best friends that he himself cited and a set of 10 best friends obtained using the proposed model. Average overlapping of these two sets was 70.9%. Additionally, the overlap level of these two sets for different demographic groups was analyzed.
Social networks based on ICT (Information and Communication Technology) are nowadays one of the most popular services based on the Internet infrastructure. They are global phenomenon that greatly affects the modern way of life. Contrary to the widespread opinion, which assumes that social networks are interesting only for private users, these networks can produce added value in companies as well. The corporate social network is a system based on the web technologies that enables agile collaboration and information exchange within company. According to the method of making connections, social networks can be divided into two groups: implicit and explicit networks. While in explicit social networks a person herself defines another person to connect with, implicit networking is determined by a person's interests and by the level of communication and collaboration with other persons. This paper addresses the question of developing implicit corporate social networks. Based on the analysed communication performed through several communication channels (i.e. e-mail, file transfer, telephone calls and instant messaging) between the employees of a multinational company, we propose an algorithm for building the social network graph. Our algorithm calculates the level of connection between employees based on the level of communication between them. We verified the proposed algorithm on our prototype test application, the FER CSN Analysis.