
Machine learning is implemented extensively in various applications. The machine learning algorithms teach computers to do what comes naturally to humans. The objective of this study is to do comparison on the predictive models in cyberbullying detection between the basic machine learning system and the proposed system with the involvement of feature selection technique, resampling and hyperparameter optimization by using two classifiers; Support Vector Classification Linear and Decision Tree. Corpus from ASKfm used to extract word n-grams features before implemented into eight different experiments setup. Evaluation on performance metric shows that Decision Tree gives the best performance when tested using feature selection without resampling and hyperparameter optimization involvement. This shows that the proposed system is better than the basic setting in machine learning.
The proliferation of internet newspapers making an Automatic Text Summarization is now a need to produce a summary that contains most of the important information from the original document. This study focused on the keyword extraction using Latent Dirichlet Allocation and Sentence Selection that used rule based concept approach to produce extractive summary. 100 Malay news documents covering general, sports, health and technology were collected from Utusan Online to evaluate the effectiveness of the system. This study used a single topic from LDA and top 10 words in the selected topic as the keywords. To evaluate, summary generated by the system was compared to summary generated by human expert using Precision Recall formula. The results showed the effectiveness of the summary generated by the system which is the best score 62.7 % that can help people read the Malay news documents in short time as the summary assist the readers to understand the important parts of the document without reading the whole document.
In Evidence-Based Medicine (EBM), medical literature is an essential resource used by clinicians and researchers. It contains research claims that summarize the critical findings of a study. However, research claims on the same topic can be contradicting. Given a clinical question, if two claims that answer the question have conflicting assertion values (Yes or No), they are considered contradictory claims. Hence, discovering the claim assertion value of a research claim is the key to detecting contradictory research claims in medical literature. In this study, we explored the usage of deep neural network (DNN) to recognize contradictory research claims in medical literature. The DNN model should determine the assertion value of a research claim against its clinical question. The model was evaluated using a publicly available corpus containing contradictory research claims from 24 systematic reviews on cardiovascular topic. Different DNN techniques such as the Global Vectors for Word Representation (GLoVe), bidirectional Long Short-Term Memory (LSTM) and Bidirectional Encoder Representations from Transformers (BERT) were implemented in building the claim assertion model. The evaluation results suggest that the BERT model performs better than LSTM and GloVe model and outperforms the previous studies’ models.
Information retrieval systems play the critical role of meeting information needs of users. Therefore, high effectiveness is expected from these systems since users make decisions based on the retrieval results. The effectiveness of these systems varies and is only known through retrieval evaluation and the main approach of evaluating these systems is the test collection model which comprises of a corpus of documents, topics and relevance judgments. One limitation of this approach is the cost of generating the relevance judgments. A recent solution to this limitation is the prediction of performance metrics at the high evaluation depths of documents using the system scores of other performance metrics computed at the low evaluation depths. However, this solution has a drawback of the inaccurate predictions of the non-cumulative discounted gain (nDCG) and the precision performance metrics at the high evaluation depths of documents using other performance metrics computed at the low evaluation depths. Therefore, this study addresses this drawback by proposing an approach that predicts the nDCG and precision performance metrics at the high evaluation depths of documents using topic scores of other performance metrics computed at the low evaluation depths. This study has shown that the proposed approach performs better predictions of the nDCG and precision performance metrics than the existing method.
Forests are essential for the protection of biodiversity and provide essential ecosystem services to humankind. Globally, 1.6 billion people depend on forests as fuel sources, construction materials, medicine, food, and freshwater sources. However, according to the monitoring service Global Forest Watch, our world lost 12 million hectares (30 million acres) of tropical tree cover in 2018, equal to 30 football pitches a minute. Surprisingly, Malaysia was among the top six countries that year with tremendous losses. Therefore, stern action is needed to increase public knowledge of ecosystem threats, improve strategic forest management decisions, and enforce land-use policies. This study aims to map deforestation in Permanent Forest Reserve (HSK) Yong in Pahang between 2017 and 2020 using satellite images of Sentinel-1 SAR. We further automate the classification of forest and non-forest by using a small and straightforward Convolutional Neural Network (CNN) model architecture. Our model used an open-source machine learning framework, namely Orfeo ToolBox TensorFlow (OTBTF). As well as machine learning, deep learning can be applied by OTBTF without image size restrictions and is computationally efficient, regardless of hardware configuration. The methodology includes data pre-processing, RGB composition, patches sampling, TensorFlow model train, TensorFlow model serve, and verification. Results show that the CNN approach with the Sentinel-1 Synthetic Aperture Radar (SAR) was 81.57 percent of the overall accuracy and the Kappa index was 0.6313. In brief, the approach mentioned in this study provides an alternative and reasonable approach for the HSK deforestation mapping in Peninsular Malaysia.
Twitter is a well-known social networking platform where users exchange information and express opinions. Since many people's interactions now take place on social media, this medium has rapidly become a source of capturing knowledge from users. The aim of this study is to find tweets using R that are then related to Ibn Khaldun's thoughts. The analysis was carried out with a simple algorithm written in R Programming. 45 keywords based on Ibn Khaldun's thoughts were constructed as the hashtag to retrieve data from Twitter. As a result of the data extraction, 1075 public tweets were collected through the search API. The simple algorithm is capable to facilitating the process of extracting data from Twitter, such that it easily known to the public tweets that exist in a given time period and this can be used as reference material for further development process. For further work, it is possible to use text mining activities and sentiment analysis approaches, as well as explore various social media platforms using R packages for decision makings.
This research aims to design a chatbot application to teach Artificial Intelligent (AI) in Malay language. CikguAIBot offers learners the possibility to learn and interact with an agent without the need for a teacher to be around. The development of CikguAIBot is based on the RAD model with the involvement of a number of experts and real users. The main focus of this paper is on the contents and flow design of the chatbot so that the objectives of the chatbot are achieved. Results from the expert review sessions were reported and a detailed evaluation strategy with the students is also included although the evaluation session is in the future plan. This research is expected to foster the usage of chatbot technology in supporting successful learning in Malaysia's education system.
Before the modern era of personal computing, it was impossible to collect and analyze data in a wide variety of methods found on big data technology. The analytical knowledge of big data room for people to identify a new cures and better understanding a number of diseases and health care. Hence, there are a number of issues that need to be further addressed. The major focus of this study is to deliver a review of studies that utilized big data issues during pandemic COVID-19 in 2020. The study is implemented by ten phases, includes defining research questions, scope review, conduct study, extracting all papers, paper screening, relevant papers, search more specific keywords, classification scheme, data extraction and systematic map. A detailed review studies were selected from January 2020 until December 2020 based on sources from Web of Science (WoS). This study highlights three key words: big data, business intelligence and business analytics. The findings of the study show that, the field of computer science is the most frequent publication in 2020 with 14 journal publications related to the field of computer science. Nevertheless, Sustainability journal showed the highest number of publications in 2020. Followed by the journals of Future Generation Computer Science and Basic and Clinical Pharmacology & Toxicology, Applied Science-Basel, and the International Journal of Information Management.
This paper presents the development of COVID-19 Malay Corpus using the WordPress content management system (CMS). The COVID-19 Malay Corpus is an information retrieval system that collects articles related to COVID-19 from Malay online news media by using a crawler or web content extractor. This proposed system used a search engine that takes search queries to retrieves the news articles in the Malay language over the COVID-19 and related. The system can be used by Malay-speaking people to search and read documents about COVID-19 in the Malay language since most of the documents available are in English. WordPress has been used for developing this system as it has more plugins, themes, and other customizations available than the other CMS. Besides, it provides many plugins that enhance website functionality, including the Octolooks Scrapes plugin and Relevanssi plugin that are used for crawler and indexing, and the Yoast SEO plugin used for search engine optimization (SEO). This system’s collection has been tested and evaluated by COVID-19 Malay Corpus documents, stopwords list, Malay root words, terms weighting, indexing result, Malay natural language query list, relevant judgment list, relevant queries feedback, and retrieval evaluation by using precision and recall.
A variety of applications have been built in recent years with the aim to extract knowledge from Al-Quran. Current knowledge representations of Al-Quran give attention primarily on conceptual ontology models that describe the semantic relations between the Quranic concepts or entities. There seems to be minimal effort towards recognizing the semantic relations between words in Quranic text, which is relatively more complex. This paper aims to present a framework for semantic knowledge representation of Al-Quran using dependency relations between words, in an attempt to boost the retrieval accuracy for Al-Quran. The semantic analysis is performed on Quranic verses according to word dependency relations using dependency parsing. Based on parsed dependencies, a set of rules are formulated to build a semantic graph of Surah Ali Imran of Al-Quran. The efficiency of the semantic representation was tested by developing a prototype question answering system. The framework was evaluated using precision and recall, First Hit Success, First Answer Reciprocal Rank and Total Reciprocal Rank by comparing the retrieved and actual answers. The results indicate that the performance of the proposed framework using word dependencies improves the semantic representation of knowledge.
In this project, a profile matching approach is proposed to assist internship placement tasks. Students often have difficulty finding and applying for internship placements, while companies spend much time sorting applications. Computer-assisted applications help job placement and onboarding tasks by mapping the potential job seekers to the corresponding companies. However, lack of focus put on internship placement, as the student profiles are limited to academic achievement without extensive work experience. A student-industry profile matching approach is proposed to evaluate students’ profiles and map the academic criteria to potential companies offering internship placement. The fuzzy matching technique is used to model the correlation ratio between the student profiles and the industry profiles.
In this paper, we survey and classify most of the information retrieval (IR) approaches to Malay text in order to assess their benefits and limitations. We also summarized the information retrieval tools and related methods, in which ontology is a widely used tool for all countries' researchers. This research selects Malay language as the primary test collection because there are more issues in Malay languages, particularly those related to deep semantics, including the use of ontology. The traditional Malay retrieval system mostly focused on syntax extraction and keywords only. Mostly this technique will ignore the semantic element and the real meaning of query text and corpus which not fulfil the requirement of the user. Most of the previous study in information retrieval was using English and Arabic language as a test collection. Therefore, advance research is needed and it will be experimented in the future work. The finding of the paper will help other researchers discover the information and research gap regarding the Malay text.
The Digital Jahai - Malay Language Repository is a platform that was created to store and translate Jahai terms into two different languages which are English and Bahasa Melayu. The system was developed to assist the needs to translate a Jahai term using a text input. Previously, the only way to translate Jahai terms to English is by using a corpus in a paper published by Burenhalt. Until recently, no system was developed to translate the Jahai language to Bahasa Melayu. Since ethnic Jahai resides in Malaysia, it is more beneficial if we can develop our own language repository system that will translate the Jahai terms to Bahasa Melayu, and preserve the language using a sustainable format which is digitization. The proposed system is capable of identifying the Jahai terms using International Phonetic Alphabet (IPA) symbols by entering the special characters using a virtual keyboard. Several additional functions such as insert new Jahai terms and listening to Jahai term pronunciation are also provided. In the future, a mobile translation application will be developed to increase the usability of the existing system.
Blockchain technology is a way to improve data integrity and data traceability, and data manipulation. Blockchain technology can be used in the drug manufacturing business process. However, despite the existence of Good Manufacturing Practice (GMP) regulations, there are still many counterfeit drugs circulating in the community. It is a very public concern and can even cause death for people who consume fake drugs. Blockchain technology can be used as a solution to minimize the circulation of counterfeit drugs. The research was conducted using a qualitative approach, namely the User Center Design (UCD). The results of a business model created based on blockchain technology are then validated in one of the largest pharmaceutical industries in Indonesia, which consists of 5 experts from various fields. Validation is carried out by means of a Forum Group Discussion. The final result of the FGD states that the drug manufacturing process business model with blockchain technology is suitable and can be used by the drug industry. With this blockchain system, it provides excellent data integration for each transaction data on each party to be able to view and carry out transactions so that data can be traced properly and give industry owners a sense of confidence in the production section of the pharmaceutical industry.
Our planet is known as a digital earth, circulating around data. Growth in data is exponential, leading to an elevated interest in Big Data Analytics, to collect, store, process, analyze and visualize unparalleled amount of data. Modern information driven society will continue to be shaped by big data, where there will be potential to extract meaningful insights and hidden patterns impacting businesses in unforeseen measures. Most employers in Malaysia provide medical benefits which includes general medical costs to hospitalization benefits and insurance coverages; with these data and information stored by the HR (Human Resource), leading to a potential to analyze and identify patterns in historical claims - these insights would lead to improved decision making to better understand employee population health and the usage of the premium coverage. In predictive analysis, common techniques applied are Decision Tree and Regression. Therefore, the aim of this research is to propose a conceptual prediction model to better understand the patterns present in the employee healthcare data while predicting if an employee would be at any health risks to understand the population health and the usage of premium coverage provided by the employer. Additionally, to apply an ensemble method called Stacking, where multiple predictive models will be combined to perform a prediction. An ensemble model will present the opportunity to build a more robust and accurate model which could be applied across various industries instead of being industry specific.
This research aims is to develop an approach in supporting component developers measure the reusability of developed software components. The approach could be used by software developers to select the reusable components during software development. Thus, this paper presents a controlled experimental design using human subjects. The overall method includes a detail explanation of the procedures and guidelines for performing the experimental activities. The experiment consists of numerous activities: designing the experimental method and procedure, designing the experimental tasks, running the controlled experiment, collecting data, analyzing the results, and recommending improvements based on the implications of earlier studies. The findings of the study can be applied to evaluate the metrics for the Component Reusability Evaluation Approach (CREA), namely documentation, observability, customisability, and the external dependency of the component.
Speaker change detection (SCD) is the fundamental task for other speech-related tasks in processing voice recordings of meetings. In this research, the Bidirectional Long Short-Term Memory (Bi-LSTM) model is used to provide synthesized features to predict speaker change from feature representations such as Mel Frequency Cepstral Coefficients and Mel Spectrogram audio features. This research explores the impact of long speaker duration in SCD within a voice meeting recording. We proposed a method to improve an existing SCD methodology that uses local maxima as a threshold to detect speaker change segments by using derivatives of Bi-LSTM predictions. Our proposed model provides a robust pipeline that uses 2 seconds as the segmentation duration which was able to achieve at most 0.76 purity along with 0.71 coverage for the voice recordings of meetings, each with the duration of 20 minutes.
There is a lack of application that promotes melody training despite research in speech-related applications. Melody training technology allows users to do repetitive practice and provide feedback on the user’s performance. Conventional melody training requires one to have face-to-face training to master the melody from a teacher. The study focuses on people who want to learn Tarannum as a preliminary step for mastering the melodic patterns for a Tarannum. The proposed application implements a curve similarity algorithm to help the user to train their Quranic recitations, following the melody of the Tarannum. The curve similarity algorithm captures the pattern of melodic recitations and provides feedback to the user. The application compares the recorded Taranum melody with the trained patterns of the Bayati recitations and provides feedback on the accuracy, which is measured using a melody curve similarity algorithm. The application offers an experience and nurtures users to repeat recitations aiming for the highest possible accuracy. The research can achieve the highest accuracy on some segments at 100% similarity with overall accuracy at 88.7% by verse. For example, the users will do their best to get a passing grade for every training. Hence, the user will keep trying until they achieve their desired result and learn the Tarannum melody.
The Coronavirus crisis has cast a shadow over the education sector; As it pushed schools, universities, and educational institutions to close their doors, to reduce the chances of its spread. This raised great concern among those affiliated with this sector, especially students preparing to take important exams. All this pushed educational institutions to switch to E-learning, as an alternative that has long been talked about over the need to integrate it into the educational process. From this standpoint, the importance of electronic exams as an alternative to the traditional paper exams appears, there are many assessment methods in E-learning, one of these assessments is the essay. The study used a variety of similarity tests and used an Arabic data set of 40 student answers. The results indicate that the Arabic Automated Essay System will improve with Cosine similarity and Arabic WordNet. Automatic Arabic Essay Scoring using WordNet is higher in terms of precision than the Automated Arabic Essay Scoring without using WordNet dependent on the mean absolute error value and Pearson correlation. The results clearly show the Cosine similarity with Arabic WordNet has the lowest error. Cosine similarity with all stemming types has the lowest error compared with the Jaccard and Euclidean similarity. Euclidean similarity has the highest error.
Breast cancer is responsible for the death of thousands of women around the world. A cancerous tumour increases temperature in the region near to the tumour, such heating is then transferred to the skin surface. Early breast cancer detection can save many lives and reduce the cost of treatment. There are many imaging techniques to do the screening such as thermography. To achieve an early detection, Computer-Aided Detection (CAD) systems are needed to classify masses in thermogram images. In this research, we propose a hybrid methodology based on Histogram of Oriented Gradients (HOG) and Discrete Wavelet Transform (DWT) to classify temporal-based thermogram images into either normal or cancerous. In this study we propose a technique to reduce the HOG coefficient vector and extract features from this vector. Two-dimensional Discrete Wavelet Transform (DWT) is applied on the original images to obtain HH (high-high) sub band. Then, HOG is applied using the HH sub band. Support Vector Machine (SVM) binary classifier is used to classify images to either normal or abnormal using the extracted features. The obtained results showed comparable results with state-of-the-art related research.