Identifying highly potential athletes is a critical yet inherently challenging process that requires comprehensive analysis of diverse factors, including physiological attributes, demographic characteristics, and social influences. This multifaceted process requires meticulous evaluation of extensive datasets to ensure both accuracy and fairness in talent identification protocols. The complexity stems from the interconnected nature of the determinants of athletic performance, where physical capabilities intersect with psychological resilience, social support systems, and environmental factors. In recent years, machine learning (ML) algorithms gain prominence in decision-making processes, offering unprecedented opportunities to uncover subtle patterns and relationships within athlete data that might otherwise remain hidden. This study systematically benchmarks the performance of several state-of-the-art ML classifiers using a novel, self-collected dataset of athlete candidates. Furthermore, an explainable AI (XAI) technique, Shapley Additive Explanations (SHAP), is applied to interpret model decisions and provide meaningful insights into key predictive factors. Experimental results demonstrate that Gradient Boosting achieves superior predictive performance (F1) across the 10-fold sets, with a mean value of 0.46. SHAP analysis reveals the critical importance of anthropometric measurements and social group features in influencing prediction outcomes. These findings collectively underscore the substantial potential of ML to revolutionize talent identification in sports while emphasizing the importance of model interpretability in fostering trust and acceptance of AIdriven decision-making processes.
The Generalized Linear Model (GLM) is an extension of the general regression model for response variables following an exponential family distribution, including normal, binomial, Poisson, negative binomial, exponential, and gamma. If the response variable is discrete and follows a Poisson distribution, then the Poisson regression model can be used for model formation. However, in its application, overdispersion often occurs, where the variance is greater than the mean. Overdispersion in Poisson regression can occur due to a large number of observations having zero values in the response variable (excess zeros). Data experiencing overdispersion and excess zeros are more suitable for using Zero Inflated Negative Binomial (ZINB) and Hurdle Negative Binomial (HNB) regressions. In this study, the models were further developed into ZINB and HNB regression models with transformed variables to improve model performance. Real-world issues related to these methods can be encountered in mortality cases, where the data used pertain to the number of toddler deaths due to pneumonia in East Java in 2022. The results of model selection using AIC show that the ZINB regression model with transformed variables is the best model in this study, with an AIC value of 58.63682. The results of the partial significance test of parameters in the ZINB regression model with transformed variables indicate that the percentage of vitamin A supplementation (x(5)), and the percentage of exclusive breastfeeding (x(7)) significantly influence the number of toddler deaths due to pneumonia.
Urban hydrological challenges, such as flooding and water resource management, require accurate rainfall data to support sustainable development. This study investigates the use of Recurrent Neural Networks (RNN) for spatial interpolation of monthly rainfall data across 31 districts in Surabaya, Indonesia, and compares its performance with the geostatistical method Cokriging. Elevation data were incorporated as an additional variable to account for geographical variability. The dataset was divided into training (26 locations) and testing (5 locations) subsets, with testing locations treated as missing data points to simulate real-world conditions. The results show that the RNN-based interpolation method achieved progressively lower Root Mean Square Error (RMSE) values from January (48.65) to April (13.78), indicating higher accuracy compared to the Cokriging method. These findings underscore the potential of RNN in addressing data gaps and spatial variability, offering robust solutions for hydrological applications in urban environments. This approach not only supports flood risk mitigation strategies but also contributes to optimizing drainage systems and water resource planning. Further research is recommended to incorporate additional environmental variables and extend the application to broader spatial and temporal contexts.
Regression is a widely used statistical modeling method. Its applications have evolved, including functional data analysis, which offers flexibility by modeling data over specific time intervals with varying case-specific observations. Functional regression has several categories, one of which is functional prediction regression. This study uses functional prediction regression to model comprehensive climate data from the World Bank Climate Change Knowledge Portal, specifically Indonesia’s average annual temperature and rainfall from 1951 to 2016. An advantage of functional prediction regression is its ability to model based on the dataset. We compare two models: the first with a functional linear effect and the second with a linear interaction combination effect. The models are estimated using boosting and the Generalized Additive Model for Location, Scale, and Shape (GAMLSS). Results show that the functional linear model better fits Indonesia's rainfall data, yielding a smaller Akaike's Information Criterion by 170.351.
The type and intensity of exercise performed by athletes play an important role in affecting blood pressure stability, putting them at risk of developing hypertension. Hypertension, or high blood pressure, is a medical condition in which the blood pressure in the arteries rises above normal limits. Hypertension in athletes becomes an essential factor in real cases if not detected early. Therefore, this study aims to model and analyse the sociodemographic and anthropometric factors that influence the incidence of hypertension. The data used in this study are primary data from 200 athlete selection participants at the University of Surabaya and the Indonesian National Sports Committee (INSC) of East Java. This research method proposes to compare the traditional approach with machine learning to prove the accuracy comparison of the model's goodness, where both approaches are proposed by considering the novelty proposed through the machine learning approach but still maximizing the traditional approach. The proposed methods are binary logistic regression, binary logistic regression with the addition of random effects, highly randomized tree, and support vector classification. The binary logistic regression model is better than the binary logistic regression model with random effects, random trees, and support vector classification because the accuracy, sensitivity, specificity, and F1-score value (68.5%, 69%, 68%, and 68.8%) is highest than the others. Other results showed that the waist circumference variable, the father's occupation variable, and the salary variable significantly affected hypertension at the 5% significance level.
In December 2019, there was a virus outbreak caused by a virus disease with a relatively high spread in Indonesia, one of which was in East Java Province. It is proven by the number of new cases on January 15, 2021, in East Java, reaching 12818 cases. This is why researchers predict the number of positive cases of COVID-19 in East Java so that the Government can anticipate an increase in the number of COVID-19 patients. This study uses data on the addition of positive COVID-19 cases in East Java from May 16, 2020, to January 24, 2021. Because the count time series data shows overdispersion, predictions are made by modeling the COVID-19 data using the INAR( ). development model, namely Double Poisson INAR( ). Several tests were carried out with data from the Double Poisson distribution, and then the ACF and PACF plots were analyzed to find the order of INARDP. After obtaining the order, the model can be constructed and estimated using MLE. Then, the prediction of adding COVID-19 cases in East Java on January 25, 2021, obtained 949 cases with an estimated error of 13.73 percent. So, the model show that the accuracy of the forecasted value with actual value is 86.17 percent.
Dalam era teknologi yang berkembang pesat, pemanfaatan perangkat lunak seperti Microsoft Excel menjadi krusial, terutama di bidang pendidikan. Meski populer, sebagian besar guru di Kabupaten/ Kota Pasuruan belum sepenuhnya memaksimalkan Microsoft Excel untuk infografis data. Tim Pengabdian Kepada Masyarakat (PKM) mengusulkan kolaborasi dalam "Pelatihan Infografis data dengan Microsoft Excel Bagi Guru Kabupaten/Kota Pasuruan" tahun 2024, bertujuan meningkatkan pemahaman dan keterampilan guru sejalan dengan visi lembaga untuk meningkatkan mutu pendidikan. Pelatihan difokuskan pada pembuatan infografis data melalui fungsi Excel, meningkatkan keterampilan guru di Kabupaten/Kota Pasuruan dalam memanfaatkan perangkat lunak ini. Pelatihan ini telah dilaksanakan oleh tim secara luring pada hari Sabtu, 13 Juli 2024 bertempat di SMPN 1 Bangil. Pelatihan diikuti oleh 24 guru di beberapa SD di Pasuruan dan berjalan lancar. Pada saat pelatihan dilakukan penilaian dengan pretest dan posttest, berdasarkan hasil tes tersebut akan dilakukan pengujian menggunakan uji t, dengan hasil yang diperoleh adalah nilai thitung lebih besar dari ttabel dan terjadi kenaikan rata-rata nilai pretest (30,83) ke posttest (67,92) yang menunjukkan bahwa pelatihan yang dilakukan berhasil meningkatkan pemahaman dan kemampuan peserta, yang ditunjukkan oleh kenaikan nilai. Hasil survei evaluasi kegiatan menunjukkan bahwa secara keseluruhan, peserta memberikan penilaian yang sangat positif terhadap kegiatan PKM ini, dengan rata-rata poin keseluruhan sebesar 4.8. Hal tersebut menunjukkan bahwa materi dan penyampaian dalam kegiatan ini berhasil menarik minat dan memenuhi kebutuhan peserta, dengan mayoritas merasa antusias, termotivasi, dan mampu meningkatkan kemampuan mereka dalam membuat infografis data menggunakan Microsoft Excel.
This research aims to develop an analytical approach to classification statistics. The proposed approach combines machine learning with optimization. Considering the urgency of research related to exploring the best methods to apply to sports data. This study proposes a novel framework that combines the k-means clustering results with the bat algorithm to optimize performance prediction for athletes in Indonesia. The proposed method aims to explore the data by comparing the classification performance of random forests, extremely randomized trees, and support vector machines. We conducted a case study using primary data from 200 respondents at Surabaya State University and the East Java National Sports Committee. The accuracy results in this study indicate that, based on the performance evaluation metric, the best approach is random forest clustering using k-means with bat algorithm optimization, achieving 81.25% accuracy, compared with other machine learning approaches. This research contributes to the field of classification statistics by introducing a novel hybrid framework that integrates machine learning, clustering, and optimization techniques to improve predictive accuracy, particularly in sports analytics. Beyond sports science, the proposed approach can be adapted to other domains that require robust performance prediction and decision support, such as health analytics, educational assessment, and human resource selection.
The COVID-19 pandemic in Indonesia has recorded an increase in stress levels of around 75% in 2021. This problem shows the importance of further testing to classify a person’s condition based on the level of depression, anxiety, and stress (DAS). A good classification requires the latest methods in its approach. Therefore, this study aims to compare suitable machine learning approaches to predict the level of DAS. This research presents nine machine-learning approaches to classify a person’s category based on the dimensions of DAS. Logistic regression, support vector classification (SVC), k-nearest neighbors, gaussian naïve Bayesian, random forest, stochastic gradient descent, linear SVC, gradient boosting, and decision tree were applied to primary data from a survey of 344 people who completed the depression, anxiety, and stress questionnaire. This study found that the two best approaches, based on the performance of the cross-validation mean scores metric, were logistic regression and support vector classification. Both methods provided an accuracy value of >85% in the classification of each dimension. In addition, other explanations were found from a health perspective. The findings in this study can be used as reference material for dealing with stress during a pandemic or post-pandemic to achieve normal conditions.
Stocks have high-profit potential but also have high risk. Many people have ways to forecast stock prices. The Geometric Brownian Motion (GBM) method forecasts stock prices. The data used in this study are closing stock price data from July 1, 2021 to August 31, 2021 taken from Yahoo! Finance. The stocks used in this research are Bank Rakyat Indonesia (BBRI), Indofood Sukses Makmur (INDF), and Telkom Indonesia (TLKM). A strategy is carried out to improve prediction accuracy by utilising the Kalman Filter (KF). This research will compare the mean absolute percentage error (MAPE) value between GBM-KF, which was manually computed and computed using the Python library. As an example of this research, for BBRI stock, the high GBM MAPE value of 9.02% can be reduced to 3.52% with manually computed GBM-KF and 3.68% with Python library computed GBM-KF. Similarly, INDF and TLKM stocks are showing a significant reduction in MAPE values to deficient levels in some cases. The GBM-KF method employing manual computing may enhance the overall precision of stock price forecasting. Future research may enhance this study by using the GBM-KF model on alternative financial instruments, integrating supplementary market data, or evaluating its efficacy under extreme market conditions.
This study aims to examine the relationship between the angle of deviation and the average time length of a physical pendulum. This study uses primary data, namely practicum report data in a physics laboratory. Data collection was conducted under closed conditions with control variables (shaft length, number of oscillations, and radius of gyration) and dependent variables (angle of deviation and average time). The analysis technique used is a statistical approach technique, namely Latin Square Design analysis. The findings of this study indicate that the variables of deviation angle and average time on the physical pendulum have a strong and unidirectional correlation relationship. This result is evident from the value of F0 > Fcritical and the value of Pr (> P) > 0, 05(significant level) where both of these prove that there is a strong reason to reject H0 which means that there is at least 1 difference in the effect of the angle of deviation on the average time. It shows that the angle of deviation affects the average time of the physical pendulum
Kusta ialah penyakit kronis yang diakibatkan Mycobacterium leprae, yang melukai saraf tepi (fungsi sensorik, motorik, dan otonom). Perawatan yang tertunda dapat mengakibatkan kerusakan permanen dalam mata, tangan, dan kaki. Penelitian ini bertujuan mengidentifikasi faktor-faktor yang memiliki pengaruh angka positif kusta di Jawa Timur. Faktor-faktor yang dapat mempengaruhi antara lain kepadatan penduduk, jumlah desa atau kelurahan yang memiliki fasilitas kesehatan, presentase penduduk dengan keluhan kesehatan, presentase masyarakat dengan fasilitas sanitasi memadai, presentase masyarakat miskin, jumlah tenaga kesehatan, serta presentase yang memiliki asuransi kesehatan. jumlah. Persentase pekerja dan mereka yang memiliki asuransi kesehatan. Metode yang digunakan adalah metode regresi binomial negatif. Ini adalah salah satu metode yang digunakan untuk mengatasi overdispersi data dalam regresi Poisson. Data penelitian ini menggunakan data yang diperoleh dari Badan Pusat Statistik dan Publikasi Dinas Kesehatan Jawa Timur tahun 2021. Hasil penelitian menunjukkan bahwa kepadatan penduduk, presentase penduduk dengan keluhan kesehatan, dan presentase penduduk miskin merupakan faktor yang mempengaruhi signifikansi penderita kusta di Jawa Timur tahun 2021. Kata Kunci: kusta, regresi poisson, overdispersi, regresi binomial negatif.
Cancer is a disease characterized by the uncontrolled growth and spread of abnormal cells in an organ of the human body. Asia is the continent that has the most significant number of new cases of cancer, with a percentage of 49.3% of the number of cancer patients in the world. Preventive action to deal with the spread of cancer is the responsibility of the government to improve the quality of health in the country, so it is necessary to take action to prevent the spread of cancer and help archieve the Sustainable Development Goals (SDGs) at the third point in the field of health. One of them is by determining the characteristics of the cancer and clustering countries in Asia based on their characteristics. This article will discuss the clustering of countries in Asia using fuzzy clustering in the form of fuzzy k-means, fuzzy Gustafson-Kessel babushka and fuzzy k-medoids. the results obtained from the analysis show that using fuzzy k-means will have a more excellent fuzzy silhouette index value compared to fuzzy Gustafson-Kessel babushka and fuzzy k-medoids, which is 0.6313.
In this increasingly advanced era, beauty products are increasing. Not only women, men also enjoy the development of this beauty product. One beauty brand that keeps up with current developments is the MS Glow brand. The MS Glow brand not only offers types of skincare that are suitable for men and women. This brand offers skincare that is licensed by BPOM and Halal MUI. Even though this brand is well known to the general public, it is necessary to identify the factors that influence customers' decisions in purchasing MS Glow products. This can help increase the popularity of MS Glow products. The cultural factor used in this research is the attitude of wanting to own what other people have or just following trends in the surrounding area. From this factor it can be concluded that the cultural factor is a "following" trend. Lifestyle factors in this research are defined as observations/interactions of customers who want to buy products. In this modern lifestyle, many customers always want products that are attractive or because of the quality of the product. According to Joesyiana, one of the most effective and efficient ways of marketing goods or services is through the word-of-mouth communication process using online media. With this marketing method, the brand owner gets the advantage that his product is better recognized by the public
Hypertension and diabetes are two medical conditions that are often associated with athletes’ health. Hypertension or high blood pressure is a condition where the blood pressure in the arteries becomes too high. Meanwhile, diabetes is a condition where the body cannot produce or use insulin properly, thereby causing high blood sugar levels. Athletes’ health is very important because they need optimal physical conditions to be able to compete effectively. Hypertension and diabetes can affect athletes’ health and their performance. Socio-demographic and anthropometric factors are believed to play an important role in the development of both conditions. The aim of this study is to determine the relationship between socio-demographic and anthropometric factors on the incidence of hypertension and diabetes in prospective athletes in athletics and determine whether prospective athletes pass the initial screening process. This study integrates bivariate logistic regression models and decision trees to analyze data collected from 200 athlete selection participants. The univariate logistic regression model showed that waist circumference, father’s occupation, and salary category 2 had a significant influence on hypertension, while BMI had a significant influence on diabetes. Meanwhile, the bivariate logistic regression model found that BMI and salary category 2 had a significant effect on hypertension. The optimal classification tree was formed using variables such as BMI, Salary Category 2, Hypertension, and Diabetes. The accuracy of the prediction data was 72%, indicating that the optimal tree is well-formed and suitable for classifying athletes’ data. This study concludes that there is a significant relationship between sociodemographic and anthropometric factors and the incidence of hypertension and diabetes in prospective athletes. This study provides valuable insight into physiological adaptation, fitness, recovery, and other factors that influence athlete performance.