In recent years, a significant portion of the content on various platforms on the internet has been found to be offensive or abusive. Abusive comment detection can go a long way in preventing internet users from facing the adverse effects of coming in contact with abusive language. This problem is particularly challenging when the comments are found in low-resource languages like Tamil or Tamil-English code-mixed text. So far, there has not been any substantial work on abusive comment detection using imbalanced datasets. Furthermore, significant work has not been performed, especially for Tamil code-mixed data, that involves analysing the dataset for classification and accordingly creating a custom vocabulary for preprocessing. This paper proposes a novel approach to classify abusive comments from an imbalanced dataset using a customised training vocabulary and a combination of statistical feature selection with language-agnostic feature selection while making use of explainable AI for feature refinement. Our model achieved an accuracy of 74% and a macro F1-score of 0.46.
Road commuting often bears unnoticed risks, contributing significantly to fatalities and injuries across the world. Alarmingly, India tops the list of harm caused by road accidents with 1.5 million annual road fatalities. In this scenario, real-time speed estimation, facilitated by vehicle-mounted cameras, can potentially be an excellent tool for averting accidents and enhancing traffic management. This paper explores the complexities of real-time speed estimation, emphasizing efficient object detection and tracking for vehicles on Indian roads through transfer learning over the base model of YOLOv8. The new model yielded an accuracy of 83 % on the Indian Vehicle Dataset. The proposed algorithm also addresses the challenge of relative speed estimation of other vehicles on the road from a vehicle mounted camera by observing the changing pixel areas of the vehicles on the video feed. Through this approach, we were able to estimate the change in the speed of vehicles satisfactorily, laying the groundwork for future advancements in practical realtime speed estimation.
Over the years, there has been a slow but steady change in the attitude of society towards different kinds of sexuality. However, on social media platforms, where people have the license to be anonymous, toxic comments targeted at homosexuals, transgenders and the LGBTQ+ community are not uncommon. Detection of homophobic comments on social media can be useful in making the internet a safer place for everyone. For this task, we used a combination of word embeddings and SVM Classifiers as well as some BERT-based transformers. We achieved a weighted F1-score of 0.93 on the English dataset, 0.75 on the Tamil dataset and 0.87 on the Tamil-English Code-Mixed dataset.
Abusive language has lately been prevalent in comments on various social media platforms. The increasing hostility observed on the internet calls for the creation of a system that can identify and flag such acerbic content, to prevent conflict and mental distress. This task becomes more challenging when low-resource languages like Tamil, as well as the often-observed Tamil-English code-mixed text, are involved. The approach used in this paper for the classification model includes different methods of feature extraction and the use of traditional classifiers. We propose a novel method of combining language-agnostic sentence embeddings with the TF-IDF vector representation that uses a curated corpus of words as vocabulary, to create a custom embedding, which is then passed to an SVM classifier. Our experimentation yielded an accuracy of 52% and an F1-score of 0.54.
As the world around us continues to become increasingly digital, it has been acknowledged that there is a growing need for emotion analysis of social media content. The task of identifying the emotion in a given text has many practical applications ranging from screening public health to business and management. In this paper, we propose a language agnostic model that focuses on emotion analysis in Tamil text. Our experiments yielded an F1-score of 0.010.
We studied the distribution of ABO blood group frequencies of the Galo and Mishing subtribes of the Adi tribal cluster in East Siang District, Arunachal Pradesh, India, in order to investigate the intertribal and temporal allelic variation. Blood groups O and AB showed higher frequencies (28.4%, 27.4%) in the Galo, whereas group O (45%) was predominant in the Mishing. Allele r is significantly different in the Galo (44.6%) and Mishing (60.3%). The chi-square test indicated significant deviations from Hardy-Weinberg equilibrium. Adi tribes show high heterogeneity and indicate significant temporal variation in ABO genotype frequencies in the Galo, Mishing, and Padam, whereas the Panggi, a small isolated subtribe of Adi, show similar and stable frequencies.