Most of data over the Internet today is hosted on outsourced third-party servers which are not trusted. Sometimes data is to be distributed to other organizations or individuals for pre-agreed use. In both of these scenarios data is susceptible to malicious tampering so there is a need for some mechanism to verify database integrity. Moreover, the authentication process should be able to differentiate between valid updates and malicious modifications. In this paper, we present a novel semi-fragile watermarking scheme for relational database integrity verification. Besides detection and localization of database tampering, the proposed scheme allows modifications to the data that need periodic updates, without requiring re-watermarking. Watermark embedding is distortion free, as it is done by adjusting the text case of selected data values resulting in retention of semantic meaning of data. Additionally, group-based embedding ensures the localization of tampering up to group level. We implemented a proof of concept application of our watermarking technique. Theoretical analysis and experiments show that even a single value modification can be detected with very high probability besides detection of attacks like tuple insertion and tuple deletion.
Information overload is a common issue faced by internet users due to the huge amount of freely available content. A single search could return millions of results and it is not feasible for a user to go through each of it to find the information relevant to his/her needs. It a tedious and time-consuming task. To address this, recommendation systems have been introduced. It is a tool that provides suggestions to users based on the user’s preference of a particular item or predicted using the ratings gathered via feedback from previous users with similar taste. The emergence of various types of recommendation system has raised the question of how these systems can be evaluated in a standardized and effective manner. It is important for researchers to quantitatively compare the performance of new recommender against the existing ones to establish that the proposed solution is indeed an improvement to the current one. There is a need for researchers as well as academics to be able to compare between two different systems and establish which the better one is based on a clearly defined numerical value. Based on this numerical value, researchers will be able to decide whether the new system requires further fine tuning. Repeated evaluation can be made after each fine tuning to determine the margin of improvement that has been made. This article will propose a methodology to perform such an evaluation of recommendation systems regardless of the domain in which the recommendation system has been used.
Nowadays many IT companies are concerned about retaining their employees due to the competitiveness of employment landscape where highly skilled and quality IT knowledge workers are in huge demand. However, the growing pace of technological change and the socio-economic factors have transmuted employees in various organizations on how best to make their skills and talents proffered to be more attractive toward employers. Challenges pertaining to the turnover rate can be attributed to several factors such as failure employees to meet the company’s objectives, less motivating work environment, employees losing out to their employers, their skills are not put to best use, low wages, and many other turnover factors. On the other hand, the growing interest of machine learning (ML) lead to the prediction of the turnover factors. This research aims to explore the factors contributing to high turnover rate based on employee productivity using ML. For this purpose, dataset of 1,470 employees with 26 variables undergo feature selection processing which is to identify the correlation between employees’ productivity and turnover features which are (i) Pearson’s Square for categorical value and (ii) Correlation for numerical value have been applied to the dataset. The analysis shows higher accuracy after going through feature selection (18 variables) instead of without feature selection (26 variables). Thus, the best algorithm that can be used for the prediction of employees’ turnover is Random Forest with a model accuracy value of 87.76%, precision 0.643, recall 0.75, and F1-Score of 0.692.
Technology has become inevitable in human life, especially the growth of Internet of Things (IoT), which enables communication and interaction with various devices. However, IoT has been proven to be vulnerable to security breaches. Therefore, it is necessary to develop fool proof solutions by creating new technologies or combining existing technologies to address the security issues. Deep learning, a branch of machine learning has shown promising results in previous studies for detection of security breaches. Additionally, IoT devices generate large volumes, variety, and veracity of data. Thus, when big data technologies are incorporated, higher performance and better data handling can be achieved. Hence, we have conducted a comprehensive survey on state-of-the-art deep learning, IoT security, and big data technologies. Further, a comparative analysis and the relationship among deep learning, IoT security, and big data technologies have also been discussed. Further, we have derived a thematic taxonomy from the comparative analysis of technical studies of the three aforementioned domains. Finally, we have identified and discussed the challenges in incorporating deep learning for IoT security using big data technologies and have provided directions to future researchers on the IoT security aspects.
Open Government Data (OGD) is regarded as an organization-level innovation that works ideally in a data openness ecosystem where a government publishes data for free use and re-use by anyone without any restrictions. These days, the extant study on OGD adoption largely focused on finding the factors that influence the OGD adopter in the adoption phase. Although these studies help to clarify the adopter's decision, the stance of the adopter after accepting the OGD is very much important. Knowing the limited availability of the literature about OGD in the post-adoption phase, this study seeks to propose a research model for OGD implementation in the post-adoption phase in Malaysia's public sector. The research model is drawn from the Technology-Organization-Environment framework and innovation adoption process as the theoretical foundation, while the OGD principles are integrated as an added construct to the research model. With this research model, researchers would be able to explain the OGD adoption continuity in the post-adoption phase from the perspective of the data provider.
The paper presents a preliminary study of current progress and the issues of OGD implementation in Malaysia. With this objective, the authors attempt to identify initial factors that influence OGD implementation in the public sectors and discern how far the OGD initiative in Malaysia has grown since its inception. The authors make the highlight of the OGD implementation phase rather than adoption phase due to the research aim is to look at the OGD activities beyond adoption. Adoption phase is where the organization is in the state of deciding whether to adopt an innovation or not, while the implementation phase is the extent where the innovation is taking into actual use. Taking from the perspective of the central agency who is leading the OGD initiative, by using interview, observation, and desk research as the research approaches, the issues pertaining to OGD implementation is consolidated into the technology-organization-environment framework. The findings have indicated that data granularity, culture, policy, resources, skills, incentives, use and participation, and external pressure are the current issues transpired in the OGD implementation. These findings are contributing to the conceptual framework of authors’ future works in determining the factors influencing OGD post-adoption in the public sectors.
Off late, the ever increasing usage of a connected Internet-of-Things devices has consequently augmented the volume of real-time network data with high velocity. At the same time, threats on networks become inevitable; hence, identifying anomalies in real time network data has become crucial. To date, most of the existing anomaly detection approaches focus mainly on machine learning techniques for batch processing. Meanwhile, detection approaches which focus on the real-time analytics somehow deficient in its detection accuracy while consuming higher memory and longer execution time. As such, this paper proposes a novel framework which focuses on real-time anomaly detection based on big data technologies. In addition, this paper has also developed streaming sliding window local outlier factor coreset clustering algorithms (SSWLOFCC), which was then implemented into the framework. The proposed framework that comprises BroIDS, Flume, Kafka, Spark streaming, SparkMLlib, Matplot and HBase was evaluated to substantiate its efficacy, particularly in terms of accuracy, memory consumption, and execution time. The evaluation is done by performing critical comparative analysis using existing approaches, such as K-means, hierarchical density-based spatial clustering of applications with noise (HDBSCAN), isolation forest, spectral clustering and agglomerative clustering. Moreover, Adjusted Rand Index and memory profiler package were used for the evaluation of the proposed framework against the existing approaches. The outcome of the evaluation has substantially proven the efficacy of the proposed framework with a much higher accuracy rate of 96.51% when compared to other algorithms. Besides, the proposed framework also outperformed the existing algorithms in terms of lesser memory consumption and execution time. Ultimately the proposed solution enable analysts to precisely track and detect anomalies in real time.
The advent of connected devices and omnipresence of Internet have paved way for intruders to attack networks, which leads to cyber-attack, financial loss, information theft in healthcare, and cyber war. Hence, network security analytics has become an important area of concern and has gained intensive attention among researchers, off late, specifically in the domain of anomaly detection in network, which is considered crucial for network security. However, preliminary investigations have revealed that the existing approaches to detect anomalies in network are not effective enough, particularly to detect them in real time. The reason for the inefficacy of current approaches is mainly due the amassment of massive volumes of data though the connected devices. Therefore, it is crucial to propose a framework that effectively handles real time big data processing and detect anomalies in networks. In this regard, this paper attempts to address the issue of detecting anomalies in real time. Respectively, this paper has surveyed the state-of-the-art real-time big data processing technologies related to anomaly detection and the vital characteristics of associated machine learning algorithms. This paper begins with the explanation of essential contexts and taxonomy of real-time big data processing, anomalous detection, and machine learning algorithms, followed by the review of big data processing technologies. Finally, the identified research challenges of real-time big data processing in anomaly detection are discussed.
In recent years, environment monitoring are of greater importance towards the area of climate monitoring, analysis, agricultural productivity management, quality assurance of water, air, alongside with other potential factors that are closely connected to industrial development and convenience of living. This research is motivated by creating awareness of smart home residents on indoor air quality, as well as providing insight of carbon dioxide emissions for industries and environmental organizations. This paper proposes an efficient solution towards environment monitoring of carbon dioxide integrated with Internet of Things capability and cloud computing technology. Aforementioned techniques will deliver highly accessible and real-time data visualization which would be greatly beneficial for Smart Homes efficiency of analysis actualization and counter-measures deployment. A monitoring architecture was developed to generate, accumulate, store and visualize carbon dioxide concentration using MQ135 carbon dioxide sensor, ESP8266 Wi-Fi module, Firebase Cloud Storage Service and Android mobile application Carbon Insight for data visualization. 2880 data points in the time frame of 10 days with a 30-second interval was collected, stored and visualized with the application of this system.
Cloud computing has seen massive growth in this decade. With the rapid development of cloud networks, cloud monitoring has become essential for running cloud systems smoothly. Cloud monitoring collects monitoring metrics from the cloud's physical and virtual infrastructures. In terms of data collection, cloud monitoring can be intrusive or non-intrusive. Monitoring data collection non-intrusively from the host operating system (OS) is a challenging task. The aim of this paper was to collect monitoring data from the host OS non-intrusively and to link those data with the cloud controller for use in monitoring. Monitoring data were collected from Procfs of the host OS and that information was linked with the monitoring dashboard on the cloud controller node. The results show that the proposed solution is an efficient, lightweight, and scalable cloud monitoring framework that produces negligible overhead.
Fog computing has been emerged as a promising paradigm with the different applications from industry to academic area. In this paper, we proposed a new concept, namely fog learning with the aim of addressing the critical thinking as a demanding issue in academic area. Due to the distributed architecture of our fog learning model, it can be used anywhere, anytime, and with the minimum latency in comparison with the existing mobile and cloud learning models to foster the critical thinking. To test the effectiveness of the proposed method, we applied the Software Usability Measurement Inventory (SUMI) as well as the acceptance test. We firstly showed that more than 74% of participants suffered from lack of critical thinking. Moreover, the participants indicated that they use several critical thinking skills during information seeking process, which shows there are relationships between critical thinking skills and the information seeking processes. According to the relationships between critical thinking skills and the information seeking processes, we designed a prototype for cultivating of critical thinking on the basis of fog learning. The results of the prototype evaluation emphasized on its positive role to foster critical thinking.
Providing access to non-confidential government data to the public is one of the initiatives adopted by many governments today to embrace government transparency practices. The initiative of publishing non-confidential government data for the public to use and re-use without restrictions is known as Open Government Data (OGD). Nevertheless, after several years after its inception, the direction of OGD implementation remains uncertain. The extant literature on OGD adoption concentrates primarily on identifying factors influencing adoption decisions. Yet, studies on the underlying factors influencing OGD after the adoption phase are scarce. Based on these issues, this study investigated the post-adoption of OGD in the public sector, particularly the data provider agencies. The OGD post-adoption framework is crafted by anchoring the Technology-Organization-Environment (TOE) framework and the innovation adoption process theory. The data was collected from 266 government agencies in the Malaysian public sector. This study employed the partial least square-structural equation modeling as the statistical technique for factor analysis. The results indicate that two factors from the organizational context (top management support, organizational culture) and two from the technological context (complexity, relative advantage) have a significant contribution to the post-adoption of OGD in the public sector. The contribution of this study is threefold: theoretical, conceptual, and practical. This study contributed theoretically by introducing the post-adoption framework of OGD that comprises the acceptance, routinization, and infusion stages. As the majority of OGD adoption studies conclude their analysis at the adoption (decisions) phase, this study gives novel insight to extend the analysis into unexplored territory, specifically the post-adoption phase. Conceptually, this study presents two new factors in the environmental context to be explored in the OGD adoption study, namely, the data demand and incentives. The fact that data providers are not influenced by data requests from the agency's external environment and incentive offerings is something that needs further investigation. In practicality, the findings of this study are anticipated to assist policymakers in strategizing for long-term OGD implementation from the data provider's perspective. This effort is crucial to ensure that the OGD initiatives will be incorporated into the public sector's service thrust and become one of the digital government services provided to the citizen.
Many organizations today store their critical business information permanently in XML format. XML data can be managed using: XML-Enabled Database (XED) systems which convert and store XML files in traditional database systems; Native XML Database (NXD) systems which store XML data natively using three main storage technologies -text-based, model-based, and schema-based techniques; and Hybrid Database systems which are comprised of both XML-Enabled and Native XML database systems. NXDs are faster than other database technologies because there is no need to convert the format of the data prior to storage. No performance evaluation has been carried out to compare all three storage strategies, hence, this paper reports on the first attempt to evaluate all three storage strategies by using open source products to measure the response time taken for each of the database basic tasks such as database creation, dataset insertion, and data manipulation. The results of the evaluation show that the schema-based storage strategy: performs 3.5 times faster than the other two storage techniques in data insertion; shows very good performance in query processing on small and large datasets; performs 10.33 times faster than text-based, and 7.5 times faster than model-based storage techniques in query processing of large datasets.
Determination of bone age known as subject age estimation from clavicle bone is a critical part of forensic age estimation in criminal proceedingsespecially when evaluation is of individuals over 18 years of age. Clavicle bone is one of the long bones in a body that islast to fuse and this is the reason why it is useful for prediction of age in post-pubertal subjects. Recently, the clavicle bone and its development have been a subject of interest in medical research. This paper provides a review of the different methods for age estimation using clavicle bone.
Voluminous amounts of data have been produced, since the past decade as the miniaturization of Internet of things (IoT) devices increases. However, such data are not useful without analytic power. Numerous big data, IoT, and analytics solutions have enabled people to obtain valuable insight into large data generated by IoT devices. However, these solutions are still in their infancy, and the domain lacks a comprehensive survey. This paper investigates the state-of-the-art research efforts directed toward big IoT data analytics. The relationship between big data analytics and IoT is explained. Moreover, this paper adds value by proposing a new architecture for big IoT data analytics. Furthermore, big IoT data analytic types, methods, and technologies for big data mining are discussed. Numerous notable use cases are also presented. Several opportunities brought by data analytics in IoT paradigm are then discussed. Finally, open research challenges, such as privacy, big data mining, visualization, and integration, are presented as future research directions.
The rapid growth of emerging applications and the evolution of cloud computing technologies have significantly enhanced the capability to generate vast amounts of data. Thus, it has become a great challenge in this big data era to manage such voluminous amount of data. The recent advancements in big data techniques and technologies have enabled many enterprises to handle big data efficiently. However, these advances in techniques and technologies have not yet been studied in detail and a comprehensive survey of this domain is still lacking. With focus on big data management, this survey aims to investigate feasible techniques of managing big data by emphasizing on storage, pre-processing, processing and security. Moreover, the critical aspects of these techniques are analyzed by devising a taxonomy in order to identify the problems and proposals made to alleviate these problems. Furthermore, big data management techniques are also summarized. Finally, several future research directions are presented.
The explosive growth in volume, velocity, and diversity of data produced by mobile devices and cloud applications has contributed to the abundance of data or `big data.' Available solutions for efficient data storage and management cannot fulfill the needs of such heterogeneous data where the amount of data is continuously increasing. For efficient retrieval and management, existing indexing solutions become inefficient with the rapidly growing index size and seek time and an optimized index scheme is required for big data. Regarding real-world applications, the indexing issue with big data in cloud computing is widespread in healthcare, enterprises, scientific experiments, and social networks. To date, diverse soft computing, machine learning, and other techniques in terms of artificial intelligence have been utilized to satisfy the indexing requirements, yet in the literature, there is no reported state-of-the-art survey investigating the performance and consequences of techniques for solving indexing in big data issues as they enter cloud computing. The objective of this paper is to investigate and examine the existing indexing techniques for big data. Taxonomy of indexing techniques is developed to provide insight to enable researchers understand and select a technique as a basis to design an indexing mechanism with reduced time and space consumption for BD-MCC. In this study, 48 indexing techniques have been studied and compared based on 60 articles related to the topic. The indexing techniques' performance is analyzed based on their characteristics and big data indexing requirements. The main contribution of this study is taxonomy of categorized indexing techniques based on their method. The categories are non-artificial intelligence, artificial intelligence, and collaborative artificial intelligence indexing methods. In addition, the significance of different procedures and performance is analyzed, besides limitations of each technique. In conclusion, several key future research topics with potential to accelerate the progress and deployment of artificial intelligence-based cooperative indexing in BD-MCC are elaborated on.
Critical thinking (CT) is meta cognitive process, considerable issue and desirable outcome for higher education in the 21st century. CT is more noticeable when information users want to seek relevant information in an appropriate time and reasonable approach. Nowadays, CT is an equipment to be qualified in each dimension of life today, and also it can be an ability during the information seeking process (lSP). The aim of this paper is to find the level of CT of postgraduate students in the University of Malaya as well as investigating the Usage of CT when they seek for information. The study has adopted quantitative research design. Watson-Glaser critical thinking appraisal- UK (WGCTA-UK) edition was used to find the level of CT of postgraduate students. Moreover, another survey was prepared by the authors to find whether postgraduate students use CTin ISP or not. Printed questionnaires are distributed among postgraduate students randomly as pilot test, 45 out of 50 responded to the surveys comprehensively. This study uses the theory that a critical thinker is able to seek information more precise and accurate than person without critical thinking. The findings from the study revealed that those postgraduate students had the highest score in of assumptions and the' lowest score in inference. Furthermore, 71% of subjects are below average and average areas of CT. The result also shows that when students seek information, they use several CTskills and dispositions such as their inference, recognition of assumption, deduction and evaluation of arguments. This is the first attempt to show that postgraduate students use their critical thinking skills (CTS) and dispositions (CTD)when they seek information.