Breast cancer detection using machine learning requires precise understanding of medical images, especially for detecting boundaries between lesions and surrounding tissues. Early detection of cancer is critical since the late detection of the disease could result in fewer options for treatment. Breast ultrasound is one of the imaging methods commonly used for screening and diagnosis because of its widespread availability, lack of ionizing radiation, and capability for real-time imaging. Nevertheless, due to the presence of speckle noise and lack of image contrast, ultrasound images are not easy to interpret, and automated lesion segmentation becomes difficult. In order to overcome the mentioned difficulties, the proposed approach combines deep learning with an advanced image preprocessing for breast ultrasound images. First, the Speckle Reducing Anisotropic Diffusion (SRAD) method is performed to filter out the noise, then the Contrast Limited Adaptive Histogram Equalization (CLAHE) is performed to increase the image contrast. UNet, Attention UNet, Residual UNet, SegNet, and DeepLabV3, were individually trained on the BUSI dataset. Despite achieving notable results in segmentation, these models exhibited few drawbacks such as overfitting to training data, biased predictions, and limited adaptability to diverse patterns within medical images. Hence, to address these issues, we propose the use of ensemble learning. UNet and ResUNet emerged as a powerful combination, attaining a Mean IoU of 0.6762, Mean Dice Score of 0.8070, IoU–Dice Balance of 0.7421, and Accuracy of 0.9634, outperforming the individual models and other ensemble configurations. Further, Explainable Artificial Intelligence (XAI) techniques using Grad-CAM are incorporated to visualize the regions influencing segmentation decisions so as to improve the interpretability and clinical reliability of the proposed framework. The results indicate that combining preprocessing, model comparison, ensemble voting, and interpretability analysis can improve BUSI lesion segmentation performance.
Can we automate the prediction of word type by the mere glance of a Sanskrit raw word? Can you guess the answer? We try to answer this question head-on in this study by exploring the University of Hyderabad (UoH) Corpus, which comprises nearly 10,000 entries annotated for phonological categories. We bypass the conventional word-level features to explore character-level TF-IDF (bi-grams and tri-grams) and mBERT contextual embeddings. The most challenging issue was class imbalance, which was very tedious to tune. The rare class, i.e., words ending with vowels, posed a significant challenge that we tackled through selective oversampling. After experimenting with a wide array of classifiers, from Logistic Regression and Linear SVMs to RBF-kernel and ensemble classifiers, XGBoost emerged as the best model. With an accuracy of 93.74
Pneumonia is a prevalent illness caused by several microbiological organisms, including viruses, fungi, and bacteria. Early diagnosis is essential for a successful therapeutic approach. Computer-aided diagnosis systems assist clinicians in identifying the ailment more precisely. In this study, we use nine well-known convolutional neural network (CNN) models for the diagnosis of pneumonia, namely the CNN model 1, CNN model 2, MobileNetv2 (version 2) Model, EfficientNetB7 Model, InceptionV3 (version 3) Model, Residual Neural Network (ResNet)—101 Model, ResNet50 Model, Visual Geometry Group-19(VGG19) Model, and Xception Model. We analyze the models utilizing F1 score and accuracy. As a result, we discover that CNN Model 1 outperforms the other models, with F1 score of 91.85
Text Summarization is one of the most important tasks of natural language processing (NLP) that revolves around the extraction of key text from a larger set of text and writing an informative summary. For Hindi and other low-resource languages, the unavailability of annotated datasets together with the complexity of the language renders the task extremely challenging. This work assesses the performance of two transformer models—mT5 and IndicBART—on abstractive summarization of Hindi text based on the SAR dataset, which is regarded as a standard benchmark for this type of task. The models are fine-tuned and tested separately based on ROUGE metrics. For further improving the quality of summarization, an ensemble strategy is adopted wherein the optimal summary is chosen among the two models per sample based on ROUGE-L recall. Experimental results demonstrate that the ensemble model outperforms the individual models, achieving ROUGE-1: 0.456, ROUGE-2: 0.27, and ROUGE-L: 0.416 (F1 scores)—showing better overall precision and recall than the other models. The results from this study prove that the ensemble of multilingual transformer models, alongside lightweight post-processing, improves summarization performance in low-resource languages, which leads to more efficient and reliable multilingual NLP systems.
In recent years, speaker recognition has emerged as a pivotal research domain, particularly with the advent of Deep Neural Network (DNN)-based embeddings exhibiting remarkable results, even for short-duration audio and text-independent tasks. Speaker recognition is instrumental for identity verification and is now extensively utilized in criminal investigations, voice biometrics, and customer service sectors. Recognizing speakers irrespective of spoken content, ResNet-based architectures adeptly extract speaker embeddings by implementing residual connections in convolutional networks and standardizing residual blocks. However, their efficacy diminishes when confronted with complex input feature spaces. To mitigate these challenges, various Feature extraction variants have been investigated using customized residual learning approach. In context of speech processing tasks, Thin-ResNet, ResNet MSE, and ResNet-based TDNN have demonstrated superior performance compared to conventional speech processing methods. In this work, Res42Net model introduces extraction of Mel-filters, Sinc Band-pass filters and Bark-Frequency psychoacoustic features directly from the audio signals as input data, ensuring significant enhancement in the model’s speakers’ representation. This paper explores and implements Res42Net model trained on three cepstral features using an open-source dataset of Hindi speakers. The proposed system is evaluated and compared against the Res2Net-based system and EER reports the superior performance of the proposed system over the existing models.
A chronic retinal disorder called glaucoma damages the optic nerve and induces blindness. The eye is one of the body’s greatest significance organs for gaining an understanding of the outside world. It is possible to use artificial intelligence (AI) to guarantee early disease diagnosis and suitable therapy. In this work, we have applied four deep learning methods: GoogleNet, AlexNet, Visual Geometry Group (VGG16), and Residual Network (ResNet50). For performance comparison, four metrics have been used, i.e., Accuracy, F1 Score, Recall, and Precision. Our study suggests that among all the classifiers, GoogleNet is showing the best performance with accuracy 89.28
Machine translation is crucial to sharing knowledge using one language to any other language. It facilitates in communicating breaking language barrier. With the advent of neural networks, Machine translation has improved with good quality translations achieving human-like translation. Availability of high quality dataset remains a challenge. Since most of the Machine Translation work has been done on high resource languages like English and Chinese, very less dataset is available for Machine translation in Indian Languages. This paper tries to present an extensive list of existing dataset available for machine translation in Indian languages. These dataset are mostly pivoted to English language or not open accessible or include limited languages. We have discussed the importance of creating an open-source dataset that contains parallel sentences for multilingual translation models that includes underrepresented languages.
The globe is more technologically, socially, and culturally united when languages are used. Interlanguage information translation is essential for the interchange of ideas and information because different native speakers speak different languages. Even though Sanskrit is an old Indo-European language, it still requires a lot of information processing work to be thoroughly studied and utilized to open up new possibilities in computer science and computational languages. This paper describes a machine translation system that can convert Sanskrit to Hindi. The presented method uses a multi-head attention mechanism to train a transformer-based neural machine translation system. This is a novel strategy that can be applied to any low-resource language with rich morphology. It is a multi-field, universal system with minimal need for human interaction. Additionally, we constructed a parallel corpus for the language pair Sanskrit and Hindi. The corpus now contains more than fifty thousand parallel sentences. The BLEU performance metric was used to automatically assess the system. With an automatic assessment metrics-derived BLEU Score of 67.8
Sanskrit is a culturally rich but low resource language. Presently, it is under represented and there is no digitized corpus available for Sanskrit-Hindi language pair publicly. We present a dataset, SHiTraD that consists of 45,500 parallel Sanskrit-Hindi sentences, collected from various sources online or offline through web or scanned from books. The source-target paired sentences were created using manual translations. This dataset is then used to train models and evaluate the translations and the results are shown using BLEU score evaluation metric comparing three architecture models 1. Encoder decoder with attention mechanism, 2. Self attention based transformer mechanism and 3. Multi-head based transformer mechanism.
This study explores neural sequence-to-sequence models with attention for automatic Sandhi boundary detection. The dataset has been drawn from the University of Hyderabad Corpus and the SandhiKosh dataset and was pre-processed into SLP1 representation and thereafter used to train multiple encoder-decoder architectures. Four models were systematically compared: (i) BiLSTM encoder with LSTM decoder with dot-product attention, (ii) LSTM encoder–decoder with attention, (iii) 2-layer stacked BiGRU encoder with GRU decoder and attention, and (iv) LSTM encoder–decoder with additional attention layers. Performance was measured using sequence-level and boundary-level accuracy. Results show that the 2-layer BiGRU encoder with GRU decoder and attention achieves the best performance among all tested in this study, with 85.10% sequence-level accuracy and 91.62% boundary-level accuracy, outperforming the BiLSTM-attention model (80.25 %, 88.64 %) and the LSTM baselines (≤ 76.41 %, ≤ 84.65 %). The attention mechanism has proven to be crucial in capturing long-range dependencies within the compound words, allowing more precise boundary prediction.These findings prove that deep recurrent architectures with attention provide a robust solution for Sanskrit Sandhi splitting. This study contributes to broader efforts in computational philology even beyond linguistic applications, thereby enabling richer digital analysis of Sanskrit literature and laying a foundation for advanced NLP tasks in morphologically rich and low-resource languages.
Within the banking industry, requests for credit cards are growing tremendously, and manually reviewing each application is frequently a tiresome task that is also prone to human error. In this situation, banks and other big financial institutions can use a machine learning model to forecast whether or not to grant the customer a credit card. Banks utilize machine learning techniques to process their financial data, extract knowledge from it, and use it for risk management and decision-making. This study has created, trained, and evaluated three classification models utilizing authentic Kaggle datasets. The main research goal is to assess and contrast the models based on how accurately they project the composition of the typical class. In this work, we examine the accuracy, F1 Score, Precision, and confusion matrix of different supervised machine learning models to estimate the probability that a credit card request would be approved. After testing three classifiers, it is discovered that Random Forest outperformed Decision Tree and Logistic Regression. Random Forest's accuracy is 94.67%, precision is 0.85, recall is 0.980, and F1 Score is 2.940.
Opportunistic networks are the most used network for communication in this modern era. Many applications are provided by this network, even end-to-end connectivity and fast communications. While the attackers are advancing quickly, the opportunistic network makes users less secure. The attackers mine the data from a random node without noticing the source or destination. An optimal deep-PRoPHET is introduced for end-to-end communication using a short path and secure data transmission to secure valuable data. In this model, the kernel extreme learning machine (KELM) is used to classify nodes according to their trustiness. KELM uses the machine learning model to classify, and PRoPHET is the routing protocol used to optimize the proposed model using its unique techniques. The parameters like buffer occupancy, message live time and current hop nodes are optimized by adaptive artificial hummingbird optimization. In the result section, the proposed model is compared with existing models and proven to be superior among all compared existing methods. Parameters like average latency, delivery probability, and overhead ratio are taken from varying message generation intervals, number of nodes, and time-to-live (TTL). The proposed model acquires the values of delivery probability as 0.6628 by a varying number of nodes, 0.868 by varying buffer size, 0.6994 by varying TTL and 0.744 by varying message generation interval. The average latency of a proposed model is 183.9858 by varying TTL, 226.6054 by a varying number of nodes, 2659.74 by varying message generation interval and 93.5064 by varying buffer size. The proposed model acquires an overhead ratio of 13.0708 for varying buffer size, 12.316 for a varying number of nodes, 15.8028 by varying message generation interval and 12.3924 by varying TTL. The proposed model achieved good results among existing models.
The type of nutrition we are receiving today, collectively with our inconsistent dietary habits and schedules, are important contributors to the rising prevalence of diabetes. The main causes of diabetes are obesity, high blood sugar, and other parameters. With a focus on the Pima Indian Diabetes dataset, this research study gives a thorough investigation of predictive modeling for diabetes using machine learning approaches. We thoroughly assess different classification algorithms, such as logistic regression, K-nearest neighbors (KNN), random forest, decision trees, Naive Bayes, and support vector machine (SVM), for effectiveness in early diabetes detection with data. Best practices for feature selection, data preprocessing, and model evaluation guide our methodical approach. Results show that machine learning has the potential to improve healthcare decision-making by giving physicians trustworthy tools for identifying people at risk for diabetes. This research advances the utilization of machine learning in healthcare.
The task of identifying speakers forms a formidable field in Speaker Recognition which compares and matches the set of utterances spoken by an unknown or a known speaker with the available trained database of reference templates. And Deep Learning has proven to deliver better prediction and computational efficiency in many application domains of Speaker Recognition. Deep Neural Networks provide the best results as a backbone of deep learning due to their portability, versatility, and computational efficiency. This work presents a Fast Fourier Transform (FFT) based extraction of MFCC and GFCC features from a set of text-independent audios from the different number of speakers containing high ground noise. A Feed-Forward Back-propagation Neural Network (FFBNN) is used to categorize the voices of the selected speakers into tensor labels during the learning phase. These tensor labels are further tested with sample voice sets to identify the correct speaker. A comparative predictive outcome showed that FFBNN worked efficiently for 50 speakers, generating 85
Every day, the internet is inundated with massive amounts of data. Unlike texts, such as academic reports and news that ordinarily originate from a single source and have a well-organized structure, dialogues involve the dynamic exchange of information between two or more speakers. The objective of a discussion may shift during the track of the conversation, and important information is dispersed across multiple speakers’, making abstractive summarization of dialogues demanding. Numerous summarizing approaches have been suggested and implemented for English and other foreign languages. On the other hand, the process of summarizing dialogue in Indian languages is still in its newborn stages. By providing insights into the performance of these models and how they adapt to a wide variety of linguistic nuances, the aim of this study is to shed light on the efficacy of these models. This work addresses the challenge of abstractive summarization of dialogues in three prominent Indian languages: Hindi, Marathi and Bengali by evaluating the effectiveness of two specific multilingual models, namely mT5-small and indicBART. According to the findings of the comparative analysis, the mT5-small model has greater accuracy in the three chosen languages.
In India, there are a large number of native languages spoken with multiple dialects, allowing the speakers to share various common and distinct pronunciation characteristics. These characteristics are derived from region to region and are highly influenced by their native languages. ‘Indian English’ is a widespread and well-known language for its many eccentricities. It is widely spoken as a second language comprising 6.8
The human voice, a dynamic signal, conveys valuable information for speaker identification, encompassing gender, age, emotions, and language. In the biometrics industry, identifying voices in real-time amidst diverse accents, tones, and noisy backgrounds is a challenging task. Voice biometry, a complex aspect of speaker identification, is gaining importance in various applications, such as user authentication, attendance systems, forensics, and banking operations, as it eliminates the need for traditional credentials like cards or passwords. Recent advancements in Human–Computer Interaction technology have made conversational tasks technically feasible. Deep Neural Learning approaches, especially Convolutional Deep Neural Networks (CDNN), have emerged as a powerful tool in the field of speech processing, surpassing traditional Speaker Identification methods. This paper introduces a novel approach using 1-Dimensional Convolutional Residual Blocks for audio classification and Speaker Identification, specifically focusing on speaker recognition from spoken Hindi language. The proposed Residual architecture significantly enhances speaker identification, even in low Signal Noise Ratio environments, achieving an impressive accuracy rate of 86.02%. This outperforms traditional Gaussian Mixture Model (GMM) and Feed Forward Back-propagation Network (FFBN) model for the same set of speakers. Future research directions may explore the classification of audio and speaker identification using various acoustic features derived from speech signals.
The textual or display-based control paradigm in human–computer interaction (HCI) has changed in favor of more natural control modalities like voice and gesture. Speech, in particular, contains a significant deal of information, revealing the speaker's inner state and intention. While word analysis makes understanding the speaker's request possible, other speech aspects reveal the speaker's attitude, goal, and motivation. As a result, it is now crucial for modern human–computer interface systems to recognize emotions from speech. Numerous techniques for sound analysis have been created in the past. This work aims to detect human emotions from their voice snippet; for this, an English language open source dataset Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) and Hindi-language dataset IITKGP-SEHSC are used. RAVDESS contains over 2000 voice samples recorded by 24 actors covering eight emotions: anger, fear, neutral, calmness, happiness, sadness, disgust, and surprise. The proposed model uses ADAM optimized deep learning model along with MFCC, chroma, and Mel band spectral energy features (MBSE) to classify and recognize eight different human vocal emotions. A multilayer perceptron (MLP) classifier is used for classification. The efficiency of the proposed model was compared to another state of the art, and the outcomes were assessed. Using the proposed structure of the model on the RAVDESS and IITKGP-SEHSC datasets, an overall accuracy of 85.19
Advanced Neural Networks are widely used to recognize multi-modal conversational speech with significant improvements in accuracy automatically. Significantly, Convolutional Neural sheets have retreated cutting-edge performance in Automatic Voice Recognition (AVR) recently more appropriately in English; however, the Hindi language has not been explored and examined well on AVR systems. The work in this article has exposed a three-layered two-dimensional Sequential Convolutional neural architecture. The Sequential Conv2D is an end-to-end system that can instantaneously exploit speech signal spectral and temporal structures. The network has been trained and tested on different cepstral features such as Frequency and Time variant-Mel-Filters, Gamma-tone Filter Cepstral Quantities, Bark-Filter band Coefficients, and Spectrogram features of speech structures. The experiment was performed on two low-resourced speech command datasets; Hindi with 27,145 Speech Keywords developed by TIFR and 23,664 (1-s utterances) of English speech commands by Google TensorFlow and AIY English Speech Commands. The experimental outcome showed that the model achieves significant performance of Convolutional layers trained on spectrograms with 91.60% accuracy, compared to that achieved in other cepstral feature labels for English speech. However, the model achieved an accuracy of 69.65% for Hindi audio words in which bark-frequency cepstral coefficients features outperformed spectrogram features.