The lack of grammar-checking tools for Tagalog and the Bikol language contributes to the digital divide, restricting educational access, digital inclusion, and the preservation of linguistic heritage. This challenge is most visible in spoken contexts, where informal speech and regional variation complicate analysis. We examine Context-Free Grammars (CFGs) and Attribute Grammars (AGs) in modeling constraints for Tagalog and the Bikol language, through three case studies: (i) R-D alternation in reported speech, (ii) ng versus nang in habitual action, and (iii) the object-focused future tense in the Bikol language verbs. While CFGs can capture these rules, they demand extensive enumeration, limiting efficiency and generalizability. AGs attach conditions directly to rules, reducing redundancy and capturing context-sensitive constraints more effectively. Implemented in LanguageTool and tested on 555 phrases, the average accuracy of AGs is at least 5% higher than CFGs and REs. We highlight AGs' potential for low-resource Philippine languages, advancing inclusive language technologies and digital equity.
This paper presents an educational initiative that integrates rule-based pattern matching into a graduate-level course on Automata and Formal Languages. The course tasked students with developing practical applications using LanguageTool, an open-source rule-based engine, to address real-world challenges aligned with the United Nations Sustainable Development Goals (SDGs). Students leveraged concepts such as regular expressions, finite-state machines, and grammars to create systems for grammar and style checking in Filipino, Cebuano, Hiligaynon, and Japanese; code syntax validation; DNA/RNA sequence pattern recognition; legal text simplification; and fake news detection. By aligning computational theory with socially meaningful applications—ranging from education and health to information integrity—the project fostered both technical proficiency and civic-mindedness. Findings suggest that rule-based NLP serves as an effective medium for reinforcing formal language concepts while also encouraging socially responsible computing. The approach demonstrates how theoretical coursework can be enriched through applied, SDG-oriented student research.
•Officially: Republic of the Philippines •7,641 islands and a population of 115 million (2022 census) •Joined O-COCOSDA under the convenorship of Prof. Chiu-Yu Tseng •Hosted O-COCOSDA in 2019 under the convenorship of Prof. Satoshi Nakamura •Ethnologue country profile •The literacy rate is 98%, 518,000 deaf •There are 186 languages: 184 living and 2 extinct (“Agta, Dicamay” in the 1960s; “Agta, Villa Viciosa” in the 1990s) •184 living languages: 175 are indigenous and 9 are non-indigenous
Studies have shown how social networking sites have been used in the political landscape as a tool to disseminate information, influence people in their political views and voting decisions, and even predict election results. This study analyzes voter preferences and identifies the topics of discussion on 2022 election-related tweets using sentiment analysis and topic modelling. Naive Bayes and Support Vector Machine are used for the sentiment analysis classifier models and Biterm Topic Modeling for identifying the most discussed topics. The results of sentiment analysis show that the Naive Bayes classifier gained a higher accuracy score of 73% than Support Vector Machine with 69%. By focusing on the leading presidential candidates, the sentiment classification revealed that Leni Robredo obtained higher positive sentiment rating than Bongbong Marcos, and is the most tweeted candidate. Significant issues regarding the candidates and the elections are determined from the topic models.
The paper contributed to the literature of COVID-19 in the Philippines by conducting an interdisciplinary study on the discourse of the Department of Health during the early phase of the pandemic in the country. The frame of Critical Discourse Analysis (CDA), the process of Human Language Technology (HLT), and the concept of Crisis Management (CM) were used in analyzing the press conferences and virtual pressers. There were 56 videos selected from 30 January to 30 May 2020 using the technique of Browsing, Linking, Shortlisting, and Focusing (BLSF). The results of the study showed that: (1) Filipino language was the medium of messaging at the height of Enhanced Community Quarantine (ECQ); (2) spectrogram images revealed a monotonic voice, which was criticized in social media; (3) theme of positivity was prevalent in describing the responses of government during a crisis; (4) discourse strategies like excessive usage of inclusive pronoun [natin], retracting the official statement, downplaying the significant COVID-related updates, and blaming the ordinary citizens were utilized to handle the deficiencies of government in managing the early phase of the pandemic. For further studies, the paper can be extended with a larger dataset.
In this paper, we contribute to social media analytics literature by incorporating user feedback towards improving Tweet classification of code-switch data. We integrate this technology in a crisis information dashboard system to consolidate significant information. The instantaneous nature of data obtained from social media makes it an ideal medium in emergency situations. Using a multiclass SVM with categories (1) Announcement, (2) Casualty and Damage, and (3) Call for Help, our test case involving typhoon Hagupit with a total of 1690 tweets resulted with an accuracy rate of 63.238% as baseline. In a simulated deployment, 67 mislabeled tweets were corrected by the users, which increased the accuracy by 1%. Future work on this study can include increasing the added instances to observe a more significant difference in metrics, and to compare the difference if only corrected mislabeled tweets were added in each iteration of retraining. Multilabel classification can also be considered.
The interaction of Filipinos transitioned to a virtual setting making social media, like Twitter, their source of information since the pandemic started. The infodemic it caused has opened up avenues to understand the characteristics of misinformation tweets regarding COVID-19. In this paper, we present the classification and analysis of misinformation tweets related to COVID-19 towards identifying themes. We used pointwise KL divergence in scoring "informativeness" and "phraseness" to extract misinformation tweets and BTM for topic modeling. With a testbed of 7,711 tweets, the classifier model identified 3,533 misinformation tweets with an accuracy of 74.25%. The results of the topic modeling were analyzed and clustered to expose possible narratives in the data set. The three narratives showed that most Filipinos use Twitter to share jokes, spread information and awareness about the virus, express opinions about the government’s response, and share tips to prevent the disease. A wider date coverage could be included in future works.
This article consists of a collection of slides from the author's conference presentation. The complete oral presentation in text form was not made available for publication as part of the conference proceedings.
This article consists only of a collection of slides from the author's conference presentation.
Brands are shifting to digital services to cater to their customers who have been spending more time online. The technology that exists today enhances customer experience and actualizes customer expectations through virtual service agents or “e-service agents” during real-time interactions. Brands in most countries have to deal with bilingual customers as globalization occurs. Business process outsourcing is among the Philippines' top foreign exchange earner aside from overseas workers' remittances. These Philippine companies offer customer service but are mostly left manned-needing constant supervision. As a solution, the researchers present a bilingual retail chatbot that could handle the two official languages of the Philippines, Filipino-based on Tagalog-and English, and their code-switching variant Taglish. The proposed bilingual retail chatbot uses k-fold grid search cross-validation on a dataset constructed by a bilingual automatic corpus engine and a combination of both (1) support vector classifier-for intent identification, and (2) hash set containment-for attribute identification.
Drafting a case decision is an intensive task for lower court judges which may be a factor in the case backlogs problem of the Philippine judiciary. The search for similar Philippine Supreme Court case decisions is a common task done manually in the trial setting in order to support the decision of the judge. Doc2Vec, a common document embedding technique in NLP, and cosine similarity are implemented in order to automatically retrieve semantically similar case decisions. The model shows to have an accuracy of 80% and exhibits a strong positive correlation with the similarity scores of a legal domain expert.
This article consists only of a collection of slides from the author's conference presentation.
Cyberbullying is a form of harassment that takes place in the internet where a bully sends a harsh message to harass the receiver. In this study, a learning model is developed using Convolutional Neural Network (CNN), which is usually used for image, and is then used to create a system for detecting cyberbullying in online game chat logs. Chat logs collected from Dota and Ragnarok were preprocessed and annotated. 60% of the data were used for training, 30% were used for testing the generated model and the remaining 10% were used for validating the developed model. After validation, the accuracy of the CNN model yielded to 99.93%. Based on the results of validation, the CNN model tends to overfit, even with regularizations applied, and was not able to generalize well. For comparison, another model was created using Naive Bayes and the accuracy yielded to 92.23%. It can be concluded that detecting cyberbullying using Naive Bayes is already possible, however, the accuracy is still not comparable to existing DNN models. While the use of CNN model results to overfitting, it is recommended to explore on other DNN architectures for online games chat logs.
We present Malasakit 2.0 (meaning "sincere care" in Filipino), an inclusive, multilingual participatory online platform with feature phone integration for collecting and analyzing quantitative and qualitative textual and audio data on disaster risk reduction (DRR) strategies. Malasakit 2.0 introduces interactive voice response (IVR) to support collection of audio data via feature phone. Malasakit utilizes peer-to-peer collaborative evaluation to identify and prioritize local DRR strategies. We present results from four field tests where 261 participants provided 1,582 evaluations of current DRR strategies, and over 950 peer-to-peer evaluations on 280 textual and audio suggestions for how local government (i. e., barangays) could better support vulnerable groups (e. g., elderly, women, children, and people with disabilities) during typhoons and floods. Results suggest that individuals who engage in disaster drills are also likely to participate in their barangay's clean-up drives to reduce flooding risk by clearing drainage pathways and that those who participate in disaster drills are also likely to have enough emergency supplies for a disaster. High-rated suggestions for DRR strategies for vulnerable groups emphasize the need for communities to establish response teams that prioritize reaching out to vulnerable groups for coordination during a disaster. Malasakit can be accessed at tiny. cc/Malasakit2.
In this paper, we present our work on isolated digit speech recognition: by classifying spectrogram images and for use in a disaster preparedness participatory toolkit. To achieve higher inclusivity, we included a voice component for a wider coverage of respondents especially those who have low literacy and those vision impaired individuals. Our methodology is through speech recognition which is a deviation from usual approaches which normally work on acoustic coefficients and features. As our initial test bed, we focused on the Filipino language — a member of the Malayo-Polynesian language family and is the national language in the Philippines. Our data covers 4,297 utterances of the Filipino digits 0 to 9 collected from 262 speakers, and divided the data into 3 parts: 70% for training, 20% for testing, and 10% for validation. We applied short-time Fourier transform on our training data and we used convolution neural networks in MatLab to classify the spectrogram images. The lowest accuracy rate during our tests is 93.02%. Analyses of the results show that background noises are the cause of the misclassified utterances which will further discussed on this paper. While the results are promising, the work can be extended to include closely related languages.