In this paper, we propose script identification of Indian language document images using MobileNetV3. Script identification at the page level is an essential task for Optical Character Recognition (OCR). OCR is a time-consuming task with a lot of computation required. To make an accurate, fast, and efficient script identification module, we have employed MobileNetV3, a mobile and CPU-friendly Convolutional Neural Network (CNN) algorithm. MobileNetV3 results are comparable to ResNet50 on the ImageNet dataset. The methodology addresses challenges in real-world scenarios, enhancing robustness. The paper demonstrates the classification of six Indian scripts (Bangla, Gurumukhi, Hindi, Kannada, Malayalam, and Tamil) and English. Four experiments have been performed. The first experiment uses MobileNetV3 with center loss (CL) and without skew detection (no angle). The second experiment uses MobileNetV3 without CL and skew detection (no angle). The third experiment is done by modifying MobileNetV3 with CL and without skew. The fourth experiment uses MobileNetV3 with CL and skew angle detection (range −10 to 10 degrees). The accuracy of the fourth experiment is 96% with skew detection. The best experiment is the first, which uses CL and cross-entropy loss and achieves an impressive accuracy of 98%.
In this research, we propose to recognize scanned line images of printed text in 11 Indian languages using CMViT, a single visual model created by merging ConvMixer with modified attention in Vision Transformer (SVTR). Recognizing the Indian language is challenging due to many reasons like cursive, making new shapes after joining in case of syllables, etc. There is also a similarity in some characters in Indian language scripts. Optical Character Recognition (OCR) is the science of recognizing and converting printed line images into analyzable, editable, and searchable forms. OCR has many use cases related to text extraction from scanned documents, handwritten documents, and scene text images. The research has been going on for the last decade by using artificial intelligence/machine learning tools to automatically analyze printed documents by converting them to editable digital text. Initially, AI-based OCR development started from RNN, LSTM, BiLSTM, and encoder-decoder architecture. The OCR is developed for 11 Indian languages named Bangla, Gujarati, Gurumukhi, Hindi, Kannada, Malayalam, Marathi, Oriya, Tamil, Telugu, Urdu, and 2 bilinguals named Hindi-English, and Marathi-English. The OCR CER for each language is in the range of 0.15% to 1.95% using the proposed network.
This paper introduces an advanced web application designed to significantly improve document digitization workflows through concurrent Optical Character Recognition (OCR) processing and intuitive user interaction. All of these are seamlessly integrated within a modular system architecture that incorporates multiple microservices. The application uniquely incorporates batch OCR processing with real-time editing capabilities, thumbnail-based navigation, sentence splitting for enhanced text clarity, and comprehensive language translation functionalities via NMT. The Character Error Rate (CER) achieved by the OCR for each language (Bangla, Gurumukhi, Hindi, Kannada, Malayalam, Marathi, and Tamil) falls within the range of 0.15% to 1.95%. The BLEU score for NMT ranges from 17.88 to 35.44 on the FLORES benchmark set. This paper explains the architecture, workflow, and performance of the web application, highlighting its contributions to document management and accessibility.
A fast-expanding subject called "Smart Face Recognition" uses IoT and machine learning to reliably identify people based on their facial traits.Security, retail, and healthcare are just a few of the sectors that are using this technology to boost customer satisfaction and boost productivity.IoT and machine learning work together to enable the collection of enormous volumes of data from numerous sources, including cameras and sensors, and utilize this data to train algorithms that can precisely identify people in real-time.Due to its precision, speed, and scalability-essential characteristics for applications like security and access control-this technology is gaining popularity.The capability of this technology to learn and adapt over time is one of its main advantages.The algorithms improve in accuracy and can better recognise people even in difficult situations, including poor light or partial obstruction, as more data is gathered and analyzed.Smart facial recognition technology has the ability to transform a number of industries and improve the user experience by delivering individualized services, safe and effective access management, and real-time monitoring.We can anticipate seeing this technology used in many brand-new and fascinating applications as it Develops.
The Marwari language is home to a vast collection of historical handwritten documents that require preservation and digitalization to deepen our understanding of Marwari history. Unfortunately, the digitization of Marwari language documents has seen limited progress thus far. To address this, deep learning techniques offer valuable solutions for text generation and expediting the digitization process. This study pioneers the recognition of printed line images in the Marwari and tackles three major challenges in Marwari OCR. Firstly, unlike Hindi, the Marwari language lacks distinct word boundaries, resulting in longer interconnected words compared to other Indian languages. Moreover, handwritten form is the predominant format for Marwari documents. Lastly, the presence of plain documents adds to the complexity of line segmentation, as the text is often written without any specific formatting or structure. In this paper, we propose a novel approach utilizing DNN to overcome the first challenge. Due to the absence of an actual annotated dataset, 74,339 synthetic Marwari line images are generated with augmentation using in-house developed Bikaner font. Our method harnesses the CNN-Transformer architecture, chosen for its flexibility in hyperparameter tuning, for text recognition. The model takes Marwari line images as input, utilizing a CNN to extract visual features, which are then fed to the transformer decoder for text generation. We evaluate our approach using an encoder based on ResNet18 and a transformer decoder, achieving an impressive accuracy of 99.53% on the test set. The primary contribution of this work lies in demonstrating the successful recognition of extremely lengthy interconnected words which will benefit the research community.
A large number of Indian documents are handwritten and India is a diverse nation with many languages. These handwritten documents contain important historical and cultural information which needs to be preserved by converting to digital format. The major problem is everyone has unique handwriting with different styles of writing. To address this problem, we have trained Handwritten Optical Character Recognition (HOCR) in eight Indian languages i.e. Bangla, Gujarati, Gurumukhi, Hindi, Kannada, Odia, Telugu, and Urdu. The datasets IIIT-HW-Dev and IIIT-HW-Telugu refer to a Devanagari dataset and a Telugu dataset respectively. The IIITINDIC-HW-WORDS consists of 872K handwritten words written in 8 Indic scripts by 135 writers. Devanagari and Telugu datasets are comprised of 95K and 120K handwritten words respectively. Tamil and Malayalam languages are excluded due to issues in the IIIT-INDIC-HW-WORDS dataset. The paper describes how the CNN-Transformer architecture leverages visual and textual features to perform OCR tasks in different languages. The model takes word images as input then CNN generates visual features, and feeds them to the transformer decoder for text generation. An encoder ResNet 18 and a decoder from a transformer have been used for all eight languages to evaluate the performance of this architecture. This architecture performed best in Kannada with just a 1.5% character error rate.
In this paper, we propose a new architecture CMViT which is formed by combining ConvMixer and modified attention in Vision Transformer (SVTR), to recognize scanned Hindi line images of printed text. The performance of the proposed model is better than other state-of-the-art models on line-level text recognition. This model is customized for text recognition tasks as the height dimension is fixed and width is variable, as all of the text images are horizontal. The ConvMixer is used before the main block of SVTR as this helps in enhancing the information captured by patches. Then these feature-rich patches are given to SVTR blocks where global and local attention focus on the text and non-text components. This block is repeated three times. After each block, the height dimension is reduced until it’s one before passing it to the CTC head for sequence alignment. The total number of Hindi training line images including real and synthetic is 782751. Using the proposed CMViT Hindi OCR, errors at character, word, and line levels are 1.21%, 3.72%, and 25.7% respectively.
Identity authentication is much needed and required in this digital age where the information can be utilized in many areas like banking, finance, insurance, education, etc. The long time in the manual authentication process is tiresome for both sides due to the exchange of data. The challenge lies in verification and information extraction from the ID card during the authentication process. There is an AI-based solution needed to reduce the authentication time. This paper aims to solve this problem by doing real-time authentication of identity cards like PAN and UIDAI using AI techniques with good accuracy. Real-time authentication is done by text detection and text recognition. The text detection is done using a differentiable binarization algorithm. We do not have a real annotated dataset for an ID number. We generated approximately 90000 identity number images synthetically with noise and blur using two fonts. This dataset is divided into training, validation, and testing sets. We present a neural encoder-decoder model with attention for converting ID number line images into editable text. Our method is evaluated based on the text output of the line image. An attention-based approach can tackle this problem in a better way in comparison to other neural techniques using CTC-based models. This paper describes the usage of OpenNMT architecture for recognition due to the flexibility of hyperparameter tuning. We evaluated the text recognition performance on scanned as well as the camera-captured identity card number images. We also compared the current recognition performance of OpenNMT with Tesseract (LSTM) on the same testbed containing 36000 images containing ID numbers only. The proposed approach outperformed the Tesseract in ID number recognition.
There have been several Optical Character Recognition (OCR) related works happened for Indian languages like Hindi, Marathi, Bangla, etc. But there is very little OCR-related work done for the Sanskrit language of Devanagari script. Sanskrit is a very complex language. The large word length and old degraded documents add more challenges to Sanskrit OCR research. Due to these challenges, the word accuracy of available OCR systems is not very high for such documents. Most of the work happened to recognize Sanskrit character recognition only. There is only one attempt to recognize the whole Sanskrit line for 10 fonts.This paper shows the study of different hyperparameters of OpenNMT architecture for Sanskrit OCR of synthetically generated color line images. A neural encoder-decoder model with attention is presented to converting line images into editable text. An attention-based approach can tackle this problem in a better way in comparison to other neural techniques using CTCbased models. The main aim of this paper is to give a detailed analysis of data preparation and various hyperparameters (like the number of LSTM layers, LSTM direction, size of character embedding vector, batch size, number of iteration, and hidden unit size) of encoder-decoder in OpenNMT, and accuracy of various combinations. This paper also concludes the best accuracy model for Sanskrit OCR using OpenNMT. The text recognition performance of the proposed method on the test set is achieved 99.44%. Our major contribution is to show text recognization of degraded line images with a variety of fonts using OpenNMT architecture. Our contribution helps the researcher community in deciding hyperparameters of encoder-decoder architecture for Sanskrit language OCR.
A large number of historical documents exist in the Modi script. These documents need to be preserved and digitized so that people know more about history written in Modi. Very little work has been done for the digitization of the Modi script documents. Deep learning-based solutions help to generate text output and reduce the digitization time for Modi documents. Most of the work happened only at the character level in the Modi OCR (Optical Character Recognition). This is the first effort to recognize the line images of the Modi script.Modi OCR has three major challenges. First, there are no word boundaries in Modi script like the Marathi language. All the words are connected so word length is more as compared to other Indian languages. Second, Modi documents are mainly handwritten. Third, Text is written on plain documents which make line segmentation more difficult.This paper aims to solve the first problem with the deep neural network. We do not have a real annotated dataset for the Modi script. We generate 303571 Modi line images synthetically with noise and blur for this experiment using in-house developed MODI_SAHO font. This paper describes the usage of OpenNMT architecture for recognition due to the flexibility of hyperparameter tuning. We use a neural encoder-decoder model with the attention mechanism for the Modi OCR. We evaluated the text recognition performance on synthetically generated line images. Our method is evaluated on the basis of the text output of line images. The accuracy of the test set is achieved 99.97%. Our major contribution is to show text recognization of very long words with the encoder-decoder architecture. Our contribution helps the researcher community in building long connected words OCR.
In this paper, we evaluated the recognition performance of CRNN (Convolutional Recurrent Neural Network) on Indian language printed documents. We compared the current CRNN performance with MDLSTM (2 dimensional LSTM) and Tesseract (LSTM) on same test bed. The CRNN outperformed the previous approaches. Several experiments are done on 7 Indian languages i.e. Hindi, Marathi, Tamil, Kannada, Malayalam, Bangla and Gurumukhi. This CRNN architecture is feed with pre-segmented lines. Dataset used contained approximately 5000 pages for each language which were then divided into training, validation and testing set. The proposed fused network i.e. CRNN, does the task of feature extraction and sequence labeling as a single unit. Wherein, Convolutional Neural Network (CNN) is pitched in especially to eliminate dependency of hand crafted features. It directly processes 2D image (binarized) of each line and extracts features based on the learning kernels. Then, Recurrent Neural Network (RNN) with BLSTM (bi directional LSTM) architecture is applied on learned features to produce inference/ sequence probabilities. The output of CTC layer transcripts these probabilities into Unicode range of the evaluated language and one blank label. These layers and their parameters are empirically selected and kept same for all the languages.
Today, the digital photos especially identity images are used in almost every form filling application. As user is uploading his/her photographs, the image quality issues will come up. Especially, if some automatic processing like face recognition is happening at back end. The image quality parameters include resolution/dimensions, size and blur. Apart form blur, rest can be checked by simple conditions. But, blur detection cannot be solved in a trivial way. Here we are proposing a deep learning based approach for detecting the blur in an image. We have prepared data comprising 250,000 identity images. Our aim is to provide a comprehensive list of clear and blur images before learning happens. Apart from naturally blurred images, we have added five types of artificial blur into the images for better generalization. We have done comparative analysis of our approach to statistical feature extractor i.e. BRISQUE which was trained on SVM. We have empirically selected the convolutional network with 4 layers. Performance evaluation shows the proposed approach is giving 98.05% on 113,000 images, and thus outperforming BRISQUE using SVM.
In this paper, we evaluated the recognition performance of BLSTM (Bidirectional LSTM) and MDLSTM (two-dimensional LSTM) neural network architecture on printed documents. We also compare the performance of 2 architectures with tesseract on same test bed. We demonstrate our experimentation on 7 Indian languages i.e. Hindi, Marathi, Tamil, Kannada, Malayalam, Bangla and Gurumukhi. The input to both the architecture will be segmented lines. The data-set used contains approximate 5000 pages for each language which then divided into train, validation and test set. The Histogram of Gradients are extracted at line level to feed into the BLSTM network. Whereas MDLSTM processes 2D image (raw pixels) of each line. The level and number of hidden layers in both the architectures are empirically selected and kept same for all the languages. The output CTC layer will contain the number of unicode present in the evaluated languages and one blank label. The input layer was fully connected to hidden layers, and these were fully connected to themselves and to the output layer. The validated result shows MDLSTM outperforms both BLSTM and tesseract for all the languages included in our experimentation.
The problem of automatic table detection has always been a great topic of debate in the field of Document Analysis and Recognition (DAR). Digital documents are efficient than their printed counterparts for storage, maintenance and republishing. Being a non-textual object of a document, tables prevent OCR system to digitize a document perfectly and distorts layout and structure of digitized documents. There is no available algorithm or method which solves this problem for all possible types of tables. This paper tackles the problem of table detection and retention by proposing a bi-modular approach based on structural information of tables. This structural information includes bounding lines, row/column separators and space between columns. Through analysis of these properties, our experiments on a dataset of above 600 images consisting of more than 829 tables have detected 90% of the table correctly.
This paper presents an unconventional method of segmenting nontext objects directly from a grayscale document image by making use of connected operators, combined with the Otsu thresholding method, where the connected operators are realized as a maxtree. The maxtree structure is used for a simplification of the image and at a later stage, it is used as a structure from which to extract the desired objects. The proposed solution is aimed at segmenting halftone images, tables, line drawings, and graphs. The solution has been evaluated and some of its shortcomings are highlighted. To the best of our knowledge, this is the first attempt at using connected operators for page segmentation.
We present a document binarization scheme that is intended at consistently binarizing a range of degraded color document images. The proposed solution makes use of a mean-shift algorithm based segmentation applied at different scales of the image and a contrast enhanced version of the popular Niblack's thresholding method. The solution has been evaluated using standard metrics used in a prominent binarization competition and has also been subject to an end-to-end evaluation by use in an OCR system. The proposed solution was found to perform at par or better than existing state of the art binarization solutions and was found to always be more consistent in performance than the state of the art.
A methodology to evaluate the text level performance of Indic OCR is proposed in this paper. Indian language OCR has been an active area of research for a long time, there has been lot of work in this challenging area. Because of the complexity of Indic scripts, newer algorithms and methodologies to improve the accuracies of recognition is constantly under development. At present, the text evaluation of all Indic OCR's is reported at character level using edit distance, since there is no accurate method to evaluate Indic OCR's for word level performance. In this paper, we describe the limitation of edit distance calculations for Indic scripts; we also propose a simple methodology to evaluate word level performance of Indic scripts considering split & merge within words; the evaluation results obtained by using this method for the evaluation of Indian language OCR's consisting of 7 languages namely Hindi, Bengali, Punjabi, Tamil, Telugu, Malayalam & Kannada is also given.