This paper examines the performance of topic modelling techniques on a newly compiled dataset of Malaysian business news, aiming to address the challenge of extracting meaningful insights from unstructured data. We compare traditional methods, such as Latent Dirichlet Allocation (LDA) and Non-Negative Matrix Factorization (NMF), with emerging methods or models, including $T$ op2Vec and BERTopic. Each model is trained or fine-tuned on five yearly datasets (2019–2023) to account for temporal variations. Our goal is to identify the most effective technique for topic discovery in this domain. The dataset includes 35,667 articles collected through web scraping from New Straits Times (Malaysia) and other Malaysian news portals. The results show that NMF consistently achieves the highest coherence scores, ranging between 0.65 until 0.80, which indicate strong topic quality and interpretability. On the other hand, $T$ op2Vec performs the worst, while LDA and BERTopic produce inconsistent results. These findings provide useful guidance, particularly the strengths and limitations of each technique, for researchers and NLP practitioners applying topic modelling to similar textual corpora.
COVID-19 has profoundly impacted all countries' lives, social habits, and economies, resulting in a swift health system breakdown. This immense effect has caused worldwide research to assess its impact on various healthcare, socio-economic and demographic factors. This paper focuses on evaluating the impact of COVID-19 in Pakistan by adopting a mathematical model consisting of five population compartments: Susceptible (S), Vaccinated (V), Exposed (E), Infected (I), and Recovered (R) by utilising COVID-19 data specific to Pakistan. The primary objective is to analyse the influence of various parameters within the model. Numerical simulations were obtained using the higher-order Runge-Kutta method for dependent variables and the basic reproduction number by varying the parameters. The sensitivity analysis was then performed to assess the effect of the parameters. From the analysis, it is revealed the key parameters, including death rate, vaccination rate, and vaccine wane rate are more sensitive to the proposed SVEIR model. The simulation of basic reproduction was also carried out by observing the simultaneous effect of the five parameters, which includes the probability of susceptibility to becoming infectious per contact, isolated infectious cases, infectious period, the average number of contacts per day per case and death rate. The simulations show that the death rate produces more variations in almost all classes of the population. Vaccination rate reveals a higher number of recovered populations and reduced infected populations, and vaccine wane rate is suitable for intermediate values of the selected interval. The basic reproduction number also remains significant for the combination of probability of susceptibility to becoming infectious per contact and death rate. These insights contribute to the understanding of the sensitivity of disease dynamics under the influence of various interaction parameters.
Stingless bees are paramount in food chain as they are important pollinators of field crops. Recent studies revealed that these bees are seriously threatened by climate change and rapid urbanization across the world. It is thus important to study the relationship between the stingless bee’s diversity and the characteristics of the locations they inhibit. At the same time, clustering algorithms is a powerful machine learning approach in exploring unsupervised data. Consequently, this study aims to explore the stingless bee diversity in Malaysia through hierarchical, k-means and DBSCAN clustering. The dataset of this study consists of individual stingless bees collected from 12 locations. It comprises 14 environmental features, 3 physical characteristics, 35 species count, 12 genera counts and 3 diversity-and-abundance weights. A four-stage methodology is employed in the study. The results show that DBSCAN effectively groups data into clusters that are well-defined, but the results are less informative. In contrast, hierarchical and k-means clustering are found producing results that provide clearer insights, with hierarchical clustering delivering notably richer results.
Indigenous knowledge (IK) stands as an integral part of intangible cultural heritage, facing the impending threat of extinction amidst the dynamic shifts in the natural and social environment. This paper explores the holistic nature of knowledge, emphasizing its interconnection with humans and nature. Sarawak, as the largest state in Malaysia, with its diverse ethnic groups, serves as a rich repository of indigenous knowledge, yet grapples with challenges such as deforestation, lifestyle changes, and documentation issues. This study also highlighted how participatory design helps bridge the digital divide in projects involving indigenous communities such as the Iban, Kelabit, and Penan. It explores initiatives to co-design digital technologies to preserve indigenous knowledge. The research motivation stems from the need to extend participatory design practice to the Orang Ulu community, focusing on Indigenous Medicinal Knowledge (IMK). The paper then further delves into the current progress of the research, future plans, and key research questions. Two challenges in this research involve selecting appropriate technology for the IMK system and strategically implementing co-design to ensure its enriching and successful outcomes.
Dogs are the main source of more than 90% of human rabies infections that pose a significant threat to public health, primarily in Africa and Asia. However, it is also one of the viral diseases that can be prevented by vaccination that affects both warm-blooded animals and humans. There are two types of rabies vaccines: pre-exposure prophylaxis and postexposure prophylaxis (PEP). Mathematical models can be valuable tools for predicting and controlling the spread of rabies disease. Thus, we introduce an SEIV (Susceptible-Exposed-Infected-Vaccinated) model incorporate vaccination control strategy to examine the transmission dynamics of rabies disease in dog population. The basic reproduction number, R-0, positively invariant and attracting region, steady states, and the stability analysis of the model are investigated. We find that there are two equilibria exist in the model, i.e., disease-free and endemic equilibria. To prove the global stability of diseasefree and endemic equilibria, the theory of asymptotic autonomous system and geometric approach have been applied, respectively. Hence, we find that the disease-free and endemic equilibria are globally asymptotically stable if R-0 < 1 and , R-0< 1respectively. Numerical simulations are performed to depict the dynamics of the model. As a conclusion, we will be able to control the disease effectively if the vaccination rate is sufficiently large.
The graph-theoretic based studies employing bipartite network approach mostly focus on surveying the statistical properties of the structure and behavior of the network systems under the domain of complex network analysis. They aim to provide the big-picture-view insights of a networked system by looking into the dynamic interaction and relationship among the vertices. Nonetheless, incorporating the features of individual vertex and capturing the dynamic interaction of the heterogeneous local rules governing each of them in the studies is lacking. The methodology in achieving this could hardly be found. Consequently, this study intends to propose a methodology framework that considers the influence of heterogeneous features of each node to the overall network behavior in modeling real-world bipartite network system. The proposed framework consists of three main stages with principal processes detailed in each stage, and three libraries of techniques to guide the modeling activities. It is iterative and process-oriented in nature and allows future network expansion. Two case studies from the domain of communicable disease in epidemiology and habitat suitability in ecology employing this framework are also presented. The results obtained suggest that the methodology could serve as a generic framework in advancing the current state of the art of bipartite network approach.
The increasing demand for sustainable energy generation brings a need for tidal current energy resource exploration around the globe. Hydrodynamic modelling is an essential aspect to explore macro tidal sites. In the current research paper, a 2D hydrodynamic model is set up by utilizing the numerical application of Delft3D. The model is validated against the database results and the two macro tidal sites are identified along the coastline of Sarawak, Malaysia. The maximum available kinetic energy flux at the identified location is 0.6 kW/m2, during peak neap tide hours. This stands as a sound justification to have a detailed tidal energy assessment study in this area in future research.
The state government of Sarawak with the help of the Sarawak Disaster Management Committee (SDMC) has continuously made the updated information on the state COVID-19 situation and its ensuing control measures available to general public in the form of daily press statements.However, these statements are merely providing textual information on daily basis though the data are in fact rich in temporal and spatial properties.Since the onset of COVID-19 pandemic, spatiotemporal analysis becomes the key element to better understand the spread of COVID-19 in various spatial levels worldwide.Hence, there is an urgent need to convert this textual information into more valuable insights by applying geo-visualization techniques and geospatial statistics.The paper demonstrates the prospect of retrieving geospatial data from publicly available document to locate, map and analyze the spread of COVID-19 up to division level of Sarawak.Specifically, map visualization and geospatial statistical analysis are performed for the list of exposed locations, which are indeed locations visited by COVID-19 patients prior to being tested positive in Kuching division, using open-source geospatial software QGIS.It is found that these exposed locations concentrate on the build-up areas in the division and are in south-west to north-east direction of the center of Kuching in September and October 2021.Despite the number of exposed locations published is relatively small compared to the number of confirmed cases reported, both are nearly strongly correlated.The insights gained from such geospatial analysis may assist the local public health authorities to impose applicable disease control interventions at division level.
Background In Sarawak, 252 300 coronavirus disease 2019 (COVID-19) cases have been recorded with 1 619 fatalities in 2021, compared to only 1 117 cases in 2020. Since Sarawak is geographically separated from Peninsular Malaysia and half of its population resides in rural districts where medical resources are limited, the analysis of spatiotemporal heterogeneity of disease incidence rates and their relationship with socio-demographic factors are crucial in understanding the spread of the disease in Sarawak. Methods The spatial dependence of district-wise incidence rates is investigated using spatial autocorrelation analysis with two orders of contiguity weights for various pandemic waves. Nine determinants are chosen from 14 covariates of socio-demographic factors via elastic net regression and recursive partitioning. The relationships between incidence rates and socio-demographic factors are examined using ordinary least squares, spatial lag and spatial error models, and geographically weighted regression. Results In the first 8 months of 2021, COVID-19 severely affected Sarawak’s central region, which was followed by the southern region in the next 2 months. In the third wave, based on second-order spatial weights, the incidence rate in a district is most strongly influenced by its neighboring districts’ rate, although the variance of incidence rates is best explained by local regression coefficient estimates of socio-demographic factors in the first wave. It is discovered that the percentage of households with garbage collection facilities, population density and the proportion of male in the population are positively associated with the increase in COVID-19 incidence rates. Conclusion This research provides useful insights for the State Government and public health authorities to critically incorporate socio-demographic characteristics of local communities into evidence-based decision-making for altering disease monitoring and response plans. Policymakers can make well-informed judgments and implement targeted interventions by having an in-depth understanding of the spatial patterns and relationships between COVID-19 incidence rates and socio-demographic characteristics. This will effectively help in mitigating the spread of the disease.
Traditionally, dengue is controlled by fogging, and the prime location for the control measure is at the patient’s residence. However, when Malaysia was hit by the first wave of the Coronavirus disease (COVID-19), and the government-imposed movement control order, dengue cases have decreased by more than 30% from the previous year. This implies that residential areas may not be the prime locations for dengue-infected mosquitoes. The existing early warning system was focused on temporal prediction wherein the lack of consideration for spatial component at the microlevel and human mobility were not considered. Thus, we developed MozzHub, which is a web-based application system based on the bipartite network-based dengue model that is focused on identifying the source of dengue infection at a small spatial level (400 m) by integrating human mobility and environmental predictors. The model was earlier developed and validated; therefore, this study presents the design and implementation of the MozzHub system and the results of a preliminary pilot test and user acceptance of MozzHub in six district health offices in Malaysia. It was found that the MozzHub system is well received by the sample of end-users as it was demonstrated as a useful (77.4%), easy-to-operate system (80.6%), and has achieved adequate client satisfaction for its use (74.2%).
Haze is an atmospheric phenomenon that occurs mostly in developing countries and is caused by tiny micro-gaseous air pollutants that affect human health, the economy, and the environment. The transboundary haze, or polluted air scattered in the atmosphere from a source location, impacts the livelihoods of the human population, lasting days to weeks. To prevent massive disruption caused by haze, it is thus important to detect the probable affected locations by knowing the source location of the pollutants in the atmosphere. Based on the idea of the Epidemiological Triangle (ET), we proposed a contact triangle named Haze Contact Triangle (HCT) be formulated, representing the interdependency of the environment, source locations, and affected locations. Therefore, this paper aims to present a proof-of-concept that in detecting the affected locations, a bipartite network model can be used for this purpose. The relationship between two nodes which are the source location and the affected location of the haze was clearly defined and quantified through a weighted link. The quantification of the two nodes is based on the environmental properties of the locations so that the dispersion of haze from the source location to the affected location can be formulated. The weight of the link between the two nodes was quantified using parameters such as the high temperature, wind directions, distance between the affected location from the source location, and the rates of smouldering fire in the formulation of the Haze Contact Strength (HCS). The formulated network is then sent through a web-based searching algorithm resulting in ranked locations. Therefore, the result is an indication of a haze hotspot, which may then be used to improve the efficacy of controlling the haze situation locally.
Mathematical modeling of hand, foot, and mouth disease (HFMD) mainly focuses on compartmental modeling approaches. It classifies human population into compartments and assumes homogeneity that regards every human has equal chance of contacting other individuals in the population. However, the transmission of HFMD is complicated and dynamic with the interactions of the intertwined biomedical and social factors. Describing the disease transmission dynamic that involves high-dimensional space is mathematically challenging. The graph theoretic bipartite network modeling (BNM) approach has the potential to handle this challenge by abstracting the real-world disease transmission system and incorporating the individual features of the bipartite nodes. This study aims to seize the advantages portrayed by the BNM approach in capturing the heterogeneous features of the entities within a disease transmission system. It intends to explore adopting the BNM approach in modeling the transmission of HFMD at Kuching, Malaysia and identify the hotspot by employing the BNM approach comprising a four-stage methodology adapted from the BNM methodology framework. The bipartite HFMD contact (BHC) network is formulated with the basic building block consisting of the location and human nodes. The individual parameters of the location and human node are incorporated. The resulting BHC network formulated comprises 10 human nodes, 20 location nodes, and 23 edges. Then, six top-ranked location nodes were identified and agreed with the chosen benchmark system. The potential HFMD hotspots are thus identified by determining the location nodes ranking. The result from this study has enabled timely and effective measures and policies to be customized accordingly by the public health authorities and related policymakers.
The COVID-19 outbreak was well-controlled in the state of Sarawak, Malaysia in year 2020. A surge in positive cases started in January 2021 and affected all districts including the rural areas which have relatively limited health facilities. Hence, we investigated the spatial patterns of COVID-19 spreading at district level for the first 16 epidemiological weeks of 2021 by spatial autocorrelation analysis and spatial panel regression model. The results show that there exists weak positive spatial autocorrelation of COVID-19 confirmed cases. Having said that, the spatial cluster of high values in both weekly rate of confirmed cases and its spatial lag emerged in the center part of Sarawak in the seventh epidemiological week. Six other districts were identified as high potential for spill overing the disease to its neighbouring districts. Among the six spatial panel regression models constructed, the spatial autoregressive model which includes the spatial lag of COVID-19 confirmed cases, apart from the other two independent variables (recovered and death), is a better-fitting model. This implies that the COVID-19 spreading in the neighbouring districts has a significant effect on the rate of confirmed cases in a particular district of Sarawak.
Small and medium enterprises face the challenge of obtaining start-up fund due to the strict rules and conditions set by banks and financial institutions. The plight yields to the growth in popularity of online peer-to-peer lending platforms which are an easier way to obtain loan as they have fewer rigid rules. However, high flexibility of loan funding in peer-to-peer lending comes with high default probability of loan funded to high-risk start-ups. An efficient model for evaluating credit risk of borrowers in peer-to-peer lending platforms is important to encourage investors to fund loans and justify the rejection of unsuccessful applications to satisfy financial regulators and increase transparency. This paper presents a supervised machine learning model with logistic regression to address this issue and predicts the probability of default of a loan funded to borrowers through peer-to-peer lending platforms. In addition, factors that affect the credit levels of borrowers are identified and discussed. The research shows that the most important features that affect probability of default are debt-to-income ratio, number of mortgage account, and Fair, Isaac and Company Score.
This paper presents the predictive power analysis of the bipartite dengue contact (BDC) network model for identifying the source of dengue infection, defined as dengue hotspot. This BDC network model was earlier formulated, verified and validated using data collected in Sarawak, Malaysia. Then, a web-based BDC network system was implemented and subsequently tested by 7 other areas in Malaysia. The data collected using the system was then used to further evaluate the predictive ability of the BDC network model. The validity period of the dengue hotspots identified by the BDC network model was measured based on the accuracy of the predictive power analysis and Spearman’s Rank Correlation Coefficient (SRCC). Based on the results, using prior one-week data was sufficient to predict the dengue hotspot for the following week and subsequent two weeks. This shows that the hotspots are valid for two weeks. The accuracy for the outbreak areas is above 60%. Most of the model reported an SRCC above 0.70 which indicated a strong positive relationship between the hotspots in the targeted model and the validated model. Due to the accuracy and SRCC values obtained, it is suggested that the BDC network model can proceed further with retrospective data for other dengue outbreak areas in Malaysia and a prospective study for the areas that participated in this study.
This work aims to formalize TRIZ modelling framework and trimming techniques through mathematical notations to lay the foundations for rigorous analysis of TRIZ as a Science of Innovation. Mathematical modelling has been employed to formalize the heuristic models of trimming. A case study was presented to demonstrate the use of the proposed modelling scheme. The paper has demonstrated the correlation of TRIZ modelling framework and trimming techniques with well-established mathematical fields such as Formal Logic, Set Theory and Graph Theory. It presents initial efforts in formalizing the functional analysis and trimming techniques as a rigorous formal approach. The acceptance of systematic innovation as a scientific discipline that can be supported by knowledge systems and can be connected to mathematical models remains a dream. This work provides directions for inquiry into this non-trivial endeavour. The value of this work will see future computational models for supporting systematic innovation. The real-life use case demonstrates the powers and gaps with regards to Genrich Altshuler's modelling of product innovation using heuristics.
Public health controls the re-emergence of infectious diseases by achieving herd immunity establishment in the population. This paper aims to consolidate two factors affecting herd immunity namely individuals' trust in public health and vaccines as well as wane of immunity in individuals into one epidemiological model in maintaining herd immunity. The model formulation adopts Individual-Based Model (IBM) approach to address the heterogeneity effects of individuals' attributes. The individuals' vaccination decision is modelled with imitation dynamics based on game theory. The model simulations use the real parameters obtained for Pertussis in Malaysia. By using different probability of individuals acquire trust, the simulation results are validated with the real Pertussis prevalence where the Root-Mean-Square-Percentage-Error (RMSPE) is used as the accuracy measurement metric. Around 70 percent to 80 percent of individuals must acquire trust towards public health and vaccines so that the simulation results become more fitted to the real Pertussis prevalence. In addition, the modified herd immunity threshold formula is derived from the model and is analyzed through parameter sensitivity analysis. Comparing to the general formula of herd immunity threshold, the modified herd immunity threshold formula provides a higher threshold for the establishment of herd immunity due to additional parameters included in the formula. The disruption of herd immunity because of the wane of immunity in individual would be countered by the vaccination of individuals who trust in public health and vaccines, hence establishing the herd immunity in the population. (Abstract)
Malaysia has introduced computational thinking skills as part of a curriculum integration update to meet the global trends in 21st-century education, focusing on empowering digital literacy. Nevertheless, a preliminary investigation revealed an apparent lack of understanding of computational thinking skills in general among teachers. The study explores the feasibility of developing a localized E-learning system to train computational thinking skills among teachers. An E-learning system, termed as myCTGWBL, was developed on the basis of a newly proposed conceptual framework to present computational thinking teaching–learning repertoire to the teachers. The hypothesis is that myCTGWBL would develop teachers' computational thinking and its position in teaching–learning understanding. myCTGWBL relevance was tested through DeLone and McLean's information system and Urbach's collaboration quality construct. To determine the success factors, partial least squares structural equation modeling was used. A total of 369 teachers participated in a two-stage survey. Participants' understanding of computational thinking and perceptions were recorded at the pre- and post-intervention phases. Open-ended questions of the surveys were analyzed using a simple text analysis technique. The closed-ended questions surveys were analyzed using SPSS Statistics 22.0. A significant improvement in teachers' computational thinking teaching–learning repertoire in a relatively short period has been recorded. Teachers also demonstrated increased confidence in the future delivering computational thinking-based lessons. The E-learning conceptual framework has illustrated the predictive power between user intent, user satisfaction, and Computational thinking (CT) knowledge benefits. Results demonstrate that myCTGWBL could be used to guide future planning when establishing CT knowledge acquisition initiatives, particularly among teachers.
Leptospirosis is a zoonotic disease that is caused by the pathogen Leptospira, and it can spread indirectly or directly from infected animals to humans. According to the official statistics from the Malaysian Ministry of Health, leptospirosis outbreaks appeared to be in the most critical condition in the recent few years. The Susceptible--Infected--Recovered compartmental model and its extensions have been applied widely in disease modeling. This paper aims to present a compartmental model for leptospirosis spread in Malaysia. Using this approach, an epidemiological model is formulated for humans and vector populations. Our results indicate that the transmission rate from susceptible to infected vectors and the vector birth rate play a significant role in determining the number of infected humans. Besides, they have an impact on the duration of the outbreak as well. The simulation results have been compared with the actual data in 2017 and the analysis shows that the proposed model is able to predict the outbreak recorded in Malaysia.
Alvin W. Yeo合作论文数3