Solid waste management remains a pressing environmental concern in rapidly developing regions such as Bali, where population growth and tourism expansion intensify waste generation pressures. This study integrates a Bayesian spatio-temporal modeling framework using the Integrated Nested Laplace Approximation (INLA) to examine the dynamics of Annual Waste Generation (AWG) and Domestic Waste Generation (DWG) across districts. The INLA approach effectively captures spatial dependence and temporal variability, providing more accurate estimates than conventional regression models.Spatial diagnostic analysis using Local Moran’s I identified nineteen significant local spatial statistics (p < 0.05), revealing four key districts with distinct spatial behaviors. Badung exhibited a strong High–High cluster pattern (Ii up to 0.50922; Z.Ii ≈ 2.87), confirming high waste generation both locally and among its neighbors, driven by intensive tourism and economic activities. In contrast, Bangli displayed a Low–High pattern (Z.Ii ≈ − 2.30), suggesting a spatial outlier effect where low waste levels are adjacent to higher-producing areas. Buleleng and Karangasem showed weaker High–Low and mixed Low–Low/High–Low associations, reflecting spatial contrasts in waste management performance.The results highlight strong spatial heterogeneity and seasonal growth trends, where fluctuations in waste generation align with major tourism and cultural cycles. By combining Bayesian inference and spatial clustering diagnostics, this study provides an empirical foundation for spatially targeted policy design and supports the Sustainable Development Goals (SDG 11 and SDG 12). The findings underscore the value of spatially explicit modeling in guiding data-driven, context-specific waste management strategies in island environments.
Indonesian Sign Language consists of two systems, Sistem Isyarat Bahasa Indonesia (SIBI) and Bahasa Isyarat Indonesia (BISINDO), which serve as the primary means of communication for individuals with hearing disabilities in Indonesia. Recent advances in computer vision and deep learning have driven significant research on automatic sign language recognition using visual data. This review analyzes studies published between 2020 and 2025 to identify research trends, dataset characteristics, deep learning architectures, and reported performance in SIBI and BISINDO recognition. Based on 18 selected studies, the review reveals a clear disparity between static and dynamic recognition tasks. Static alphabet recognition using CNN-based and object detection models consistently achieves high accuracy in controlled environments. In contrast, dynamic recognition for words and continuous signing remains considerably more challenging due to temporal complexity, signer variability, and environmental factors. The findings also highlight limitations in publicly available datasets, particularly for sentence-level dynamic recognition. Overall, this review synthesizes current challenges and research opportunities, guiding the future development of robust, deployable Indonesian Sign Language recognition systems.
The government has recently adopted mobile applications to enhance service delivery for citizens. However, these applications often generate mixed reactions among users. Many citizens express their opinions through reviews and ratings on the Google Play Store, providing valuable information for sentiment analysis. Leveraging this, the present paper introduces the Indonesian Government Application Review (IGAR) dataset, a collection of 617,722 user reviews from six popular government-related applications in Indonesia: Mobile JKN, MyPertamina, KAI, JMO, Satusehat, and BMKG. The reviews, originally written in Indonesian, were manually annotated as positive, neutral, or negative based on rating scores. Among the dataset, positive sentiment accounts for 336,449 reviews, negative sentiment with 246,898 reviews, while 34,375 reviews are categorized as neutral. To extend the usability of the dataset for broader research contexts, all reviews were translated into English and further processed using the Valence Aware Dictionary and sEntiment Reasoner (VADER) for automated sentiment labeling. Through VADER classification, 324,660 reviews were identified as positive, 173,329 as neutral, and 119,733 as negative. This dataset thus provides a valuable resource for advancing sentiment classification research using machine learning and deep learning model on government-related applications in Indonesia.
The EmoTweetID dataset is a large-scale open resource of Indonesian tweets curated for emotion classification and word embedding, addressing the scarcity of publicly available datasets for this low-resource language. Tweets were collected from the X platform (formerly Twitter) using the snscrape library, yielding >4.5 million raw tweets based on basic and derived emotion keywords from Ekman’s six basic emotions. After cleaning to remove duplicates, irrelevant tweets, and sensitive elements, a corpus of 3,126,987 clean unlabeled tweets was compiled from the first scraping stage. The second stage gathered 2,243 clean tweets annotated through lexicon-based and manual methods, into six emotion classes based on Ekman’s basic emotions: anger, disgust, fear, joy, sadness, and surprise. The manual annotation by three psychology students used majority voting and achieved substantial inter-annotator agreement, with a Fleiss’ Kappa score of 0.7323. Two pretrained embeddings, Word2Vec and fastText, were trained on the corpus using a 300-dimensional skip-gram architecture to enrich semantic representation. Baseline evaluations using BiLSTM with fastText achieved a weighted F1-score of 0.8285 on the human-annotated set, demonstrating the practical utility of the dataset and embeddings for the downstream emotion classification task. Publicly available on Mendeley Data, the EmoTweetID dataset provides a valuable foundation for advancing Indonesian natural language processing by enabling unsupervised pre-training and supervised multi-class emotion classification.
Classroom engagement analysis plays an important role in understanding students' learning behaviors in a non-intrusive manner. However, many existing computer vision-based approaches rely on complex deep learning architectures and time-consuming dataset construction, which limit their applicability in practical classroom settings. This study proposes a skeletal keypoint-based pipeline for classroom engagement analysis that combines person detection and single-person pose estimation. The extracted keypoints are represented in a low-dimensional form and used as input features for lightweight engagement classification models. Two classifiers, namely Ridge Classifier and Gradient Boosting, are evaluated to assess the effectiveness of the proposed pipeline. Experimental results on a test set of 432 samples show that the Ridge Classifier achieves an accuracy of 0.90 with a macro-average recall of 0.94. In contrast, Gradient Boosting achieves an accuracy of 0.96 and a macro-average precision of 0.98. In addition, the proposed pipeline significantly improves data collection efficiency, achieving approximately a sixteen-fold reduction in pose data acquisition time compared to conventional approaches. These results demonstrate that the proposed pipeline provides a practical solution for individual-level classroom engagement analysis under controlled classroom.
The increasing prevalence of food and environmental allergies has intensified the demand for reliable computational approaches for protein allergen prediction. Recent advances in protein language models have enabled endto-end sequence representation learning without reliance on handcrafted features. In this study, ProtBERT was optimized and calibrated for protein allergen prediction using curated datasets of allergenic and non-allergenic proteins. ProtBERT embeddings were coupled with a lightweight neural classification head and optimized through learning-rate tuning, dropout regularization, partial fine-tuning of upper transformer layers, and class-balanced loss functions. Model calibration and decision-threshold optimization were further applied to enhance probabilistic reliability. Following optimization, the ProtBERT-based model achieved an accuracy of 0.9392, F1-score of 0.9339, and ROC-AUC of 0.9757, demonstrating strong discriminative capability directly from raw amino acid sequences. For contextual evaluation, ProtBERT performance was compared with physicochemical feature-based machine learning approaches, which served as reference baselines. Despite differences in modeling paradigms, ProtBERT offers advantages in scalability, reduced feature engineering, and compatibility with foundation-model frameworks. These findings indicate that optimized and calibrated ProtBERT models provide a robust sequence-based framework for protein allergen prediction and represent a promising foundation for future hybrid and large-scale computational allergology systems.
Background:Breast cancer metastasis remains a major contributor to morbidity and mortality, partly due to the disease's marked biological heterogeneity. Identification of reliable biomarkers associated with metastatic progression is essential for improving risk stratification and clinical management. Matrix metalloproteinase-2 (MMP-2) and routine hematological and metabolic parameters have been implicated in cancer progression, but their roles in breast cancer metastasis remain inconsistent across populations. Methods:This cross-sectional study evaluated general, hematological and metabolic parameters in breast cancer patients from Dr. Kariadi General Hospital, the main referral hospital for Central Java, Indonesia. Associations with metastasis were assessed using univariate analyses, followed by least absolute shrinkage and selection operator (LASSO) logistic regression for variable selection. Multicollinearity was assessed using variance inflation factor (VIF), and selected variables were entered simultaneously into a multivariable binary logistic regression model. Associations between selected hematological parameters and estrogen receptor (ER) and progesterone receptor (PR) status were also examined. Associations between clinicopathological variables and histologic grade were analyzed using the Kruskal-Wallis test. Results:Histologic grade was the only general characteristic significantly associated with metastasis (P < 0.001). In univariate analysis, metastatic patients showed significantly higher MMP-2, WBC count, and SGOT levels, whereas lower MCV and uric acid levels were detected. Multivariable logistic regression identified MMP-2 (P = 0.015; OR = 1.124, 95% CI: 1.023-1.236) and platelet count (P = 0.047; OR = 0.973, 95% CI: 0.947-1.000) as independently associated with breast cancer metastasis. Platelet count was also significantly associated with ER status (P = 0.039), but not with PR status (P = 0.054). Kruskal-Wallis analysis demonstrated significant differences in MMP-2 and uric acid levels across histologic grades I-III. Conclusion:Elevated MMP-2 levels and reduced platelets were independently associated with breast cancer metastasis, while platelets were additionally associated with ER status. These findings suggest that MMP-2 and platelet count may serve as potential biomarkers associated with metastatic behavior and hormone receptor status in breast cancer.
Computer vision is widely used for livestock farming for automated behaviour monitoring, biometric measurement, and real-time tracking. The amount of research on this topic is growing, making it important to understand the field's structure and direction through bibliometric analysis. In response, this study examines Scopus-indexed publications from 2020 to 2025 on computer vision applications for livestock farming. The dataset contained 429 documents, and using VOSviewer to analyze keyword co-occurrences, citations, and authorship to map the field's structure. Power BI was used to visualize citation distribution, publication distributions, and the most active sources. The findings show a steady rise in research output, peaking in 2024 with 110 publications, with China emerging as both the most cited country and most productive. The most active publishing source is Computers and Electronics in Agriculture. Keyword mapping showed 5 thematic clusters, emphasizing deep learning methods, real-time monitoring, animal welfare, and farm automation. The analysis also showed research gaps, including limited species diversity, dominance of CNNs and YOLO methods, underrepresentation of health monitoring applications, and a lack of standardized datasets. Overall, the study provides a thorough view of the research landscape and suggests future studies to diversify methodologies, expand species coverage, and improve dataset availability.
Physical activity is an established protective factor for colorectal cancer (CRC), but it is unclear if genetic variants modify this effect. To investigate this possibility, we conducted a genome-wide gene–physical activity interaction analysis. Using logistic regression (1-d.f), two-step screening and testing method (EDGE), and joint tests (3-d.f), we analyzed interactions between common genetic variants across the genome and physical activity in relation to CRC risk. Self-reported physical activity levels were categorized as active (≥ 8.75 MET-h/wk) vs. inactive (< 8.75 MET-h/wk; 39,992 participants) and as study- and sex-specific quartiles of activity (42,602 participants). Physical activity was inversely associated with CRC risk overall (OR [active vs. inactive] = 0.85; 95
Non-small cell lung carcinoma (NSCLC) accounts for approximately 85% of lung cancer cases, and most of them are only detected after metastasis. Given the disease complexity, virtual drug screening is essential for the identification of potential novel drugs serving as alternative treatments. A graph-based convolutional network is proposed to predict drug-target interactions (DTI) in NSCLC. The network utilizes information from bonds and atoms to construct the drug’s molecular graph, supported by molecular fingerprints to highlight chemical features in the molecule. While most DTI models solely focus on the protein sequence information, this model incorporates gene ontology as graph node features, with protein networks being connected by k-mer edges. The model outperforms conventional machine learning approaches, achieving 92.85% accuracy and an AUC of 98.26% in identifying NSCLC-related drugs. Out of 8,862 DrugBank drugs passing the drug-likeness evaluation via Lipinski’s Rule of Five, 325 drug compound candidates were identified through model screening. A random sampling of the drugs is also performed across different model evaluation threshold results by molecular docking, confirming the model’s ability to assess its drug properties. This research shows a model capable of screening potential drugs through its interaction with causative proteins, suggesting its potential as a reference for other complex diseases.
Background:This study analyzed the nongenetic risk factors that contributed to colorectal cancer (CRC) incidence in the South Sulawesi population through a case-control study. Methods:The sample consisted of 89 cases and 84 controls, aged between 19-86, with 99 males and 74 females from different ethnic groups. Univariate analysis was carried out using chi-square, Fisher's exact test, t -test, and Mann-Whitney U test. Significant nongenetic risk factors were selected through the logit model L1 regularization, adjusted for age, gender, and ethnicity. The analyzed risk factors were the patient's weight, height, body mass index (BMI), defecation location, exercise habit, smoking habit, marital status, occupation, education level, and distance to the nearest health center. The estimated odds ratio from the logit model was used to analyze the significance of the selected risk factors. Results:The significant risk factors from the logit model were smoking habit, education level, marital status, distance to the nearest health center, and weight. CRC cases were more likely to have lower education (OR = 1.819, 95% CI 1.354-2.443), residing in remote areas (OR = 1.44, 95% CI 1.17-1.772), experiencing decreasing weight (OR = 1.03, 95% CI 1.013-1.048). Controls were more likely to be non-smokers (OR = 0.325, 95% CI 0.149-0.707) and unmarried (OR = 0.161, 95% CI 0.036-0.716). Conclusion:The study determined that other nongenetic risk factors, including education level, distance to the nearest health center, weight, smoking habit, and marital status, contributed to the CRC incidence within the South Sulawesi population. The study emphasized the importance in accounting for these risk factors for further, targeted CRC preventions.
This table includes the overall sample description stratified by colorectal cancer (CRC) status and smoking status.
Plant identification is fundamental for ecological monitoring and ensuring biodiversity. Conventional identification of plant species, however, is a tedious process. Automatic plant identification using supervised learning also faces limitations in data annotation. In this study, DINOv2 selfsupervised learning is employed alongside Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction methods to examine existing plant patterns in the PlantCLEF 2025 image dataset. DINOv2 embeddings visualization using PCA depicts Bromus Sterilis L. as a vertically elongated cluster, while the other plant embeddings are forming a tight cluster. Conversely, in UMAP visualization, the plant embeddings are more wellseparated, with a higher silhouette score, lower Davies-Bouldin Index (DBI), and higher Calinski-Harabasz Index (CHI) compared to PCA (silhouette score 0.4401 vs 0.0905; DBI 2.9388 vs 3.8379; CHI 7910.6325 vs 3640.4362). Examination of UMAP visualization also uncovers subtle patterns in the PlantCLEF 2025 images, such as phenotypic phases in Hedera helix L., photograph angle variations in Scandosorbus intermedia and Aria edulis, and even mislabeled images, particularly in Scandosorbus intermedia and Aria edulis. This study has proven that unsupervised learning, embedding generation, and visualization can help the taxonomist in classifying plants from images, without the need for annotation labels.
Solar power is a feasible replacement for fossil fuel-based energy sources since it reduces carbon footprints and greenhouse gas emissions, qualities needed for the fight against climate change. Still, solar energy constantly faces societal, environmental, and technical obstacles. Advanced computing techniques such as artificial intelligence (AI) can help to alleviate some of these technical challenges. This paper uses the PRISMA approach on the Scopus database for bibliometric analysis to investigate AI applications in Solar Photovoltaics (PV). We then arrange publication themes, analyse the publishing pattern to identify the most significant journals and publications using author keyword data. Co-occurrence analysis and trend visualisation were assessed using VOSviewer and Biblioshiny. Based on the findings, India, China, the United States, and Saudi Arabia dominate the publishing output, with high citation counts and extensive collaboration networks. Interestingly, four institutions in Saudi Arabia emerged as the lead with the most publication counts in this field. To identify current research gaps and future directions, we also evaluated hotspots and recent advancements in the AI applications within Solar PV research. AI-driven forecasting, fault detection, performance monitoring, cybersecurity, and Internet of Things (IoT) integration, which enables real-time and autonomous decision making, could potentially take the front stage in future renewable energy research.
This file includes the expression imputation statistics and included SNPs from the elastin net models.
Supplementary Figure from Beyond GWAS of Colorectal Cancer: Evidence of Interaction with Alcohol Consumption and Putative Causal Variant for the 10q24.2 Region
This study investigates the application of the Variational Autoencoder (VAE) model for classifying lung images obtained from chest X-ray data, with a primary focus on utilizing the encoder component of the VAE as a visual feature extractor. The VAE architecture is implemented using six different convolutional neural network backbones, namely VGG16, VGG19, ResNet50, MobileNetV2, AlexNet, and DenseNet121, to evaluate and compare their effectiveness in learning latent representations. The resulting latent vectors are used as inputs to a Multi Layer Perceptron classifier. Experimental results show that the VAE model with the ResNet50 backbone achieves the highest classification performance, with accuracy, precision, recall, and $\mathbf{F 1}$ score all reaching 82 percent. Although the model demonstrates strong overall performance, analysis of the confusion matrix reveals that it still struggles to differentiate between normal and abnormal cases, especially in terms of sensitivity to abnormal conditions.
Major Depressive Disorder (MDD), commonly known as depression, impacts over 300 million people worldwide. Diagnosing this condition often relies on subjective judgment, emphasizing the need for an objective approach. Deep learning, particularly Recurrent Neural Networks (RNNs), offers a promising solution, especially for analyzing multivariate time series data. Researchers have explored the use of RNNs and their variants to detect depression severity. Therefore, this study conducts experimental testing to identify depression severity using first derivative techniques and feature engineering on the RNN model and its variations. The first derivative is calculated for each frame during the subject's interview, and feature engineering techniques focus on the eyes and lips, known to be associated with depression, including calculations of upper and lower lip distances, eye openness, and more. The RNN model, incorporating feature engineering, achieved the best results among the proposed methods, with a Mean Absolute Error (MAE) of 5.04 and Root Mean Squared Error (RMSE) of 6.03. However, further performance enhancement is possible by increasing the number of layers and neurons, considering the current model's relative simplicity due to limited resources.