This study describes how to develop a machine learning system for the translation of Indian languages (Hindi, Gujarati, and Punjabi) using a form of deep neural network (DNN). We are employing a form of RNN called long short-term memory (LSTM). The state-of-the-art algorithm for sequential data is recurrent neural networks (RNNs). This research work pertaining to RNN-based translation of Indian Languages involved experimenting with nmt models (unidirectional nmt and bidirectional) and their hyperparams to devise results important from real-world perspective. We employed the neural machine translation (NMT) model to get the best results. Unlike the standard phrase-based translation system, which consists of many small sub-components that are tweaked independently, this approach allows a single system to be trained directly on the source and target text and attempts at neural machine translation.
In this paper, we propose script identification of Indian language document images using MobileNetV3. Script identification at the page level is an essential task for Optical Character Recognition (OCR). OCR is a time-consuming task with a lot of computation required. To make an accurate, fast, and efficient script identification module, we have employed MobileNetV3, a mobile and CPU-friendly Convolutional Neural Network (CNN) algorithm. MobileNetV3 results are comparable to ResNet50 on the ImageNet dataset. The methodology addresses challenges in real-world scenarios, enhancing robustness. The paper demonstrates the classification of six Indian scripts (Bangla, Gurumukhi, Hindi, Kannada, Malayalam, and Tamil) and English. Four experiments have been performed. The first experiment uses MobileNetV3 with center loss (CL) and without skew detection (no angle). The second experiment uses MobileNetV3 without CL and skew detection (no angle). The third experiment is done by modifying MobileNetV3 with CL and without skew. The fourth experiment uses MobileNetV3 with CL and skew angle detection (range −10 to 10 degrees). The accuracy of the fourth experiment is 96% with skew detection. The best experiment is the first, which uses CL and cross-entropy loss and achieves an impressive accuracy of 98%.
In this research, we propose to recognize scanned line images of printed text in 11 Indian languages using CMViT, a single visual model created by merging ConvMixer with modified attention in Vision Transformer (SVTR). Recognizing the Indian language is challenging due to many reasons like cursive, making new shapes after joining in case of syllables, etc. There is also a similarity in some characters in Indian language scripts. Optical Character Recognition (OCR) is the science of recognizing and converting printed line images into analyzable, editable, and searchable forms. OCR has many use cases related to text extraction from scanned documents, handwritten documents, and scene text images. The research has been going on for the last decade by using artificial intelligence/machine learning tools to automatically analyze printed documents by converting them to editable digital text. Initially, AI-based OCR development started from RNN, LSTM, BiLSTM, and encoder-decoder architecture. The OCR is developed for 11 Indian languages named Bangla, Gujarati, Gurumukhi, Hindi, Kannada, Malayalam, Marathi, Oriya, Tamil, Telugu, Urdu, and 2 bilinguals named Hindi-English, and Marathi-English. The OCR CER for each language is in the range of 0.15% to 1.95% using the proposed network.
In the era of Artificial Intelligence (AI), significant progress has been made by enabling machines to understand and communicate in human languages. Central to this progress are parsers, which play a vital role in syntactic analysis and support various Natural language Processing (NLP) applications, including Machine Translation and sentiment analysis. This paper introduces a robust implementation of an optimized Head-Driven Parser designed to advance NLP capabilities beyond the limitations of traditional Lexicalized Tree Adjoining Grammar (L-TAG) based Parser. Traditional parser, while effective, often struggle with the capturing complexities of natural languages, especially translation between English to Indian languages. By leveraging Bi-directional approach and Head-Driven techniques, this research offers a revolutionary enhancement in parsing frameworks. This method not only improves performance in syntactic analysis but also facilitates complex tasks such as discourse analysis and semantic parsing. This research involves experimentation the Bi-Directional Parser on a dataset of 15,000 sentences, resulting a reduction in derivation variations compared to conventional TAG Parsers. This advancement highlights how Head-Driven Parsing can overcome traditional constraints and provide more reliable linguistic analysis. The paper demonstrates how this new implementation not only builds on the strengths of L-TAG but also addresses its limitations and contributes to expanding the scope of Tree Adjoining Grammarbased methodologies and advancing the field of Machine Translation.
The video camera is essential for reliable activity monitoring, and a robust analysis helps in efficient interpretation. The systematic assessment of classroom activity through videos can help understand engagement levels from the perspective of both students and teachers. This practice can also help in robot-assistive classroom monitoring in the context of human–robot interaction. Therefore, we propose a novel algorithm for student–teacher activity recognition using 3D CNN (STAR-3D). The experiment is carried out using India’s indigenously developed supercomputer PARAM Shivay by the Centre for Development of Advanced Computing (C-DAC), Pune, India, under the National Supercomputing Mission (NSM), with a peak performance of 837 TeraFlops. The EduNet dataset (registered under the trademark of the DRSTATM dataset), a self-developed video dataset for classroom activities with 20 action classes, is used to train the model. Due to the unavailability of similar datasets containing both students’ and teachers’ actions, training, testing, and validation are only carried out on the EduNet dataset with 83.5% accuracy. To the best of our knowledge, this is the first attempt to develop an end-to-end algorithm that recognises both the students’ and teachers’ activities in the classroom environment, and it mainly focuses on school levels (K-12). In addition, a comparison with other approaches in the same domain shows our work’s novelty. This novel algorithm will also influence the researcher in exploring research on the “Convergence of High-Performance Computing and Artificial Intelligence”. We also present future research directions to integrate the STAR-3D algorithm with robots for classroom monitoring.
Due to concerns over the evolving quantum computing scenario and its potential threat to existing cryptosystems, including Public-Key Infrastructure (PKI) used in electronic signatures, there is a need to enhance the eSign components with quantum-resistant algorithms. One such algorithm is CrystalsDilithium, based on lattice cryptography, which can perform PKI operations in the eSign system. In addition to being resistant to quantum attacks, Long-Term Validation (LTV) is also utilized to ensure the validity of a document for an extended period. This is achieved by referring to PAdES/CAdES profiles and a time stamping service, which provides an accurate and trustworthy record of when a particular electronic document, file, or transaction was created, modified, or sent. Given the significance of eSign as a crucial digital service worldwide, migrating to post-quantum components will strengthen the system’s security. Therefore, the proposed research provides implementation, analysis and experimentation to ensure the eSign system components are robust against quantum computing threats.
This paper presents a state-of-the-art virtual research lab (vTAG) for creating, updating, analyzing, and maintaining multilingual large-scale tree-adjoining grammar for natural languages (NL). vTAG has designed to be language independent and tested its performance by constructing grammar of many Indian and European languages with significant reductions in grammar development time by auto rule creation mechanism with the help of a supervised machine learning algorithm. vTAG provides an integrated development environment (IDE) that can generate grammar of any natural language based on tree-adjoining grammar (TAG) formalism by using a specially designed graphical user interface without focusing on the programming aspect. vTAG also contains an advanced age experimental workspace to evaluate and improvise TAG-based parsers, part-of-speech tagger, sematic text analyzer, and associated NLP tools. vTAG can be utilized as an ‘interactive training kit’ for upcoming young talent to empower them for learning and developing advanced tools solving practical problems in NLP field. Furthermore, it can serve as foundation for understanding and building end-to-end machine translation solutions. vTAG uses object oriented model (OOM) for handling the NLP resources so every entity within the environment is in form of an object that can be encrypted which makes linguistic resources convenient for maintenance and secure to exchange and distribution, thus vTAG consists of a workbench, development kit, a graphical user interface, experimental workspace, and language resources on a single platform by combining major concepts from modern computer science, artificial intelligence, and linguistics.
A large number of Indian documents are handwritten and India is a diverse nation with many languages. These handwritten documents contain important historical and cultural information which needs to be preserved by converting to digital format. The major problem is everyone has unique handwriting with different styles of writing. To address this problem, we have trained Handwritten Optical Character Recognition (HOCR) in eight Indian languages i.e. Bangla, Gujarati, Gurumukhi, Hindi, Kannada, Odia, Telugu, and Urdu. The datasets IIIT-HW-Dev and IIIT-HW-Telugu refer to a Devanagari dataset and a Telugu dataset respectively. The IIITINDIC-HW-WORDS consists of 872K handwritten words written in 8 Indic scripts by 135 writers. Devanagari and Telugu datasets are comprised of 95K and 120K handwritten words respectively. Tamil and Malayalam languages are excluded due to issues in the IIIT-INDIC-HW-WORDS dataset. The paper describes how the CNN-Transformer architecture leverages visual and textual features to perform OCR tasks in different languages. The model takes word images as input then CNN generates visual features, and feeds them to the transformer decoder for text generation. An encoder ResNet 18 and a decoder from a transformer have been used for all eight languages to evaluate the performance of this architecture. This architecture performed best in Kannada with just a 1.5% character error rate.
The Revolution of the Artificial Intelligence (AI) has started when machines could decipher enigmatic symbols concealed within messages. Subsequently, with the progress of Natural Language Processing (NLP), machines attained the capacity to understand and comprehend human language. Tree Adjoining Grammar (TAG) has become powerful grammatical formalism for processing Large-scale Grammar. However, TAG mostly rely on Grammar which is created by Languages expert and due to structural ambiguity in Natural Languages computation complexity of TAG is very high o(n^6). We observed that rules-based approach has many serious flaws, firstly, language evolves with time and it is impossible to create grammar which is extensive enough to represent every structure of language in real world. Secondly, it takes too much time and language resources to develop a practical solution. These difficulties motivated us to explore an alternative approach instead of completely rely on the rule-based method. In this paper, we proposed a Statistical Parsing algorithm for Natural Languages (NL) using TAG formalism where Parser makes crucial use of data driven model for identifying Syntactic dependencies of complex structure. We observed that using probabilistic model along with limited training data can significantly improve both the quality and performance of TAG Parser. We also demonstrate that the newer parser outperforms previous rule-based parser on given sample corpus. Our experiment for many Indian Languages, also provides further support for the claim that above mentioned approach might be an awaiting solution for problem that require rich structural analysis of corpus and constructing syntactic dependencies of any Natural Language without much depending on manual process of creating grammar for same. Finally, we present result of our on-going research where probability model will be applying to appropriate selection of adjunction of any given node of elementary trees and state chart representations are shared across derivation.
Language is the primary mode of communication. Communication is the only way to convey our thoughts and emotions to others. But there are many languages that we don't know how to speak and we can't learn all languages so quickly which is why Machine Translation systems were invented to help us communicate with anyone from anywhere. Researchers started working on Machine Translation Systems in the 1950s and have developed various amazing techniques to make our communication easy. Machine Translation Evaluation (MTE) Methodology checks the accuracy of the translations done by Machine Translation Tools. It is extremely important because while developing the translation systems, it constantly checks the performance and helps us to make appropriate changes in the system for bringing accuracy. Mostly it is done by comparing the output of machine translation systems with the translation(s) done by Human beings, there are also some techniques that do not require any reference sentences. In our project, we are using four Automated Machine Translation Evaluation metrics which are TER (Translation Error Rate), METEOR (Metric for Evaluation of Translation with Explicit Ordering), BLEU (Bilingual Evaluation Understudy), and NIST (National Institution of Standards and Technology) for English-Hindi Translation.
Speaker recognition is a prominent area of study within speech technology. In recent times, there has been a notable transition towards embedding-based end-to-end speaker recognition approaches. These methods enable speaker recognition without the need for retraining when encountering new speakers in real-world scenarios. Earlier models were trained on limited monotonous hand-crafted datasets, which worked on pattern matching, but failed to be robust because of architectural limitations and insufficient data. Research shows that training deep neural networks, which rely on attention mechanisms, on extensive datasets containing weakly labeled and diverse samples can significantly improve their resilience when applied to various classification tasks downstream. Drawing upon knowledge acquired from extensive datasets like VoxCelebl and VoxCeleb2, along with specially crafted proprietary data tailored to handle various modalities, we initiated an investigation into the capabilities of these versatile techniques for building speaker recognition models. Our attention was directed towards two reputable models, ECAPA-TDNN [1](Emphasized Channel Attention, Propagation, and Aggregation) and Titanet[2], as we aimed to comprehend the architectural enhancements and assess their impact on the model's ability to generalize effectively. Also, we propose a faster and scalable inference pipeline using Elasticsearch.
Machine translation (MT) is an important part of natural language processing (NLP). The recent advancement in deep learning (DL) has enabled us to implement the MT techniques in neural networks. Deep neural networks are powerful models that give powerful results when we train the model on a large set of data. Word embeddings play an important role in representing the syntactic and semantic relationships between the words in the form of low-dimensional dense vectors. Word2Vec or GloVe vectors can be used to initialize these word embeddings. A many-to-many encoder-decoder model helps us encode the meaning of the source sentence in a single vector to be decoded and translated into another language. In this chapter, the Recurrent Neural Network (RNN) architecture is used to implement our sequence-to-sequence model with gated RNN units to design an MT model for Indian languages like Hindi and Gujarati.
In this paper, we propose a new architecture CMViT which is formed by combining ConvMixer and modified attention in Vision Transformer (SVTR), to recognize scanned Hindi line images of printed text. The performance of the proposed model is better than other state-of-the-art models on line-level text recognition. This model is customized for text recognition tasks as the height dimension is fixed and width is variable, as all of the text images are horizontal. The ConvMixer is used before the main block of SVTR as this helps in enhancing the information captured by patches. Then these feature-rich patches are given to SVTR blocks where global and local attention focus on the text and non-text components. This block is repeated three times. After each block, the height dimension is reduced until it’s one before passing it to the CTC head for sequence alignment. The total number of Hindi training line images including real and synthetic is 782751. Using the proposed CMViT Hindi OCR, errors at character, word, and line levels are 1.21%, 3.72%, and 25.7% respectively.
Background:Machine learning (ML) prepares and trains a model through supervised or unsupervised learning methods. Sputum, a respiratory tract secretion, is a common laboratory specimen that aids in diagnosing respiratory diseases, including pulmonary tuberculosis (TB). Gram stain is an easy, cost-effective stain, which may be applied to sputum smears to screen out an unsatisfactory sample. ML model may help in screening sputum smears.Methods:This collaborative project was carried out from June 2020-July 2021. In this study, a color-based segmentation ML algorithm using K-Means clustering was developed. A library of stained sputum smears was built. The Bartletts criteria (based on neutrophil and squamous cell count) for screening and selecting satisfactory sputum smears were used. A smartphone camera was used to take several photographs of satisfactory, as well as unsatisfactory, smears. The image segmentation algorithm was applied to medical image analysis, color-segmentation of sputum images was done. The hue saturation value (HSV) color ranges were defined on a prototype image. Then, all connected pixels were identified as a single object, and morphological operations were applied.Results:Usage of AI-driven model on the slide-image revealed the slide adequacy as the cell count was acceptable based on Bartlett's criteria. Both the manual cell counts (Range: 126-203 neutrophils, 14-47 squamous cells) and the model counts (Range: 117-242 neutrophils, 14-37 squamous cells) are within acceptable limits.Conclusion:The use of a model to screen a large number of sputum slides may be a boon in resource-limited settings where trained microscopists may not be easily available.
We have unlabeled data in different ways- news article, and countless other types of documenting text. Large volumes of data are produced from many online sources such as emails, www, organization's electronic health records, and databases. These data must be classified to avoid information loss and to boost data discovery and retrieval more quickly. This research is about classifying tweets, the New York Times, or other social media. It classifies the data into predefined generic classes, such as "business," "sad," "style," "advertising," "events," "news," etc., using author information and some features. The approach is to reduce the noise and identify tweets or any data classes as it is evident that an article can fall in more than one category.
For developing and using various features provided by the internet by all the sections of the society, it is important to make the technology accessible irrespective of the language. So, the translation between the languages becomes important criterion. Previously, the machine translation has been used to translate between different languages, which eased the communication between people from different linguistic backgrounds. But in the last couple of years, the advancement and enhancement of Deep Neural Network (DNN) allowed us to rewrite our approach for machine translation completely. This paper discusses the process and the methodology to build the Machine translation system using Deep Neural Network (DNN) for translation between different Indian languages, using the combination of two ideas, i.e., Recurrent Neural Networks and Encoding.
Different types of research have been done on video data using Artificial Intelligence (AI) deep learning techniques. Most of them are behavior analysis, scene understanding, scene labeling, human activity recognition (HAR), object localization, and event recognition. Among all these, HAR is one of the challenging tasks and thrust areas of video data processing research. HAR is applicable in different areas, such as video surveillance systems, human-computer interaction, human behavior characterization, and robotics. This paper aims to present a comparative review of vision-based human activity recognition with the main focus on deep learning techniques on various benchmark video datasets comprehensively. We propose a new taxonomy for categorizing the literature as CNN and RNN-based approaches. We further divide these approaches into four sub-categories and present various methodologies with their experimental datasets and efficiency. A short comparison is also made with the handcrafted feature-based approach and its fusion with deep learning to show the evolution of HAR methods. Finally, we discuss future research directions and some open challenges on human activity recognition. The objective of this survey is to give the current progress of vision-based deep learning HAR methods with the up-to-date study of literature.
In the 21st century, natural language processing (NLP) has obtained much prominence for human–machine interaction (HMI). With this interest in natural language processing (NLP) has grown significantly, numerous NLP tools (e.g., morphology, the tagger, and a parser, etc.) have been developed all over the world. Despite having huge importance and requirements, we have noticed gaps for having a comprehensive single framework or platform, which encompass all NLP-related tools and technologies for promoting the research in NLP and sharing the knowledge and resources among NLP researchers required for understanding and building the solution for HMI. Our objective is to apply Software engineering in natural language processing with the concept of an object-oriented model by using a collection of reusable objects by defining the communication protocol, consisting of a set of rules that must be applied to exchange data between two NLP modules. We proposed state of art ivrE—A virtual environment for creating, modifying, executing, and analyzing various NLP solutions and technology. The proposed idea is broadly based on to define own ivrE-NLP object framework model that permits the developer to create, modify, and execute the application and analyze their outcomes by operations on visual representations of the modules. A variety of NLP-based applications (tools, modules, and plugins) already exist, they can publish into store available with environment so it can be used by research community at large. To develop complete NLP framework or platform, we require much more than just assembling or collecting these tools or modules at one place, no matter how good any tool or module is working individually. It requires not only the standards and a set of protocols, but also requires a compliant composition than a pre-defined algorithms and their implementation. In brief, we require a comprehensive open framework to bundle, manage, and integrate set of NLP tools, modules, components, applications, algorithms, and define their associated rules, comprehensive data structures, and knowledge.
Deep neural network architectures are highly in-corporated in modern (progressive) automatic speech recognition(ASR) systems which outperforms conventional and hybrid systems. Even after adapting to such system we still facing a problem to make balance between accuracy and model size (no. of trainable parameters). We present our experiment in building a robust encoder-decoder based end-to-end ASR model for the Hindi language. Starting with, we have trained the QuartzNet model with two different architectures QuartzNet-5x5 and QuartzNet-15x5 by keeping the same training parameters and dataset. The speech corpus consists of 2648.9 hours of Hindi labelled audio data collected from various domains and sources. We have analyzed the performance of both models and observed that QuartzNet-15x5 has a radical improvement of 24% in accuracy. For building and training the models, we have used the PARAM SIDDHI AI system and training recipes from the open source NeMo toolkit.
With the rapid progress in the technology and data in the public domain, the machine translation and data science have made remarkable progress. In this paper, we discuss our specific use case of developing machine translation system for English to Hindi and Hindi to English language translation. For this system, we have used the daily proceedings of the Lok Sabha as data and developed NMT-based machine translation system on the top of already available rule-based machine translation system. Developed system has been evaluated using bilingual evaluation understudy (BLEU) as well as the human evaluation metrics using comprehensibility and fluency. In machine translation (MT), there is the trend of measuring post-editing time, and thus, we have also evaluated our system by measuring post-editing time using open-source tool.